MEGA Hub

Beyond Global Latents: Chunk-Based Sparse Grid VAE for Scalable 3D Modeling

Authors

Do you know Kaiyi Zhang?You can claim authorship or link another user.Do you know Zhihao Liang?You can claim authorship or link another user.Do you know Haolin Liu?You can claim authorship or link another user.Do you know Qingxiang Lin?You can claim authorship or link another user.Do you know Zeqiang Lai?You can claim authorship or link another user.Do you know Yunfei Zhao?You can claim authorship or link another user.Do you know Bowen Zhang?You can claim authorship or link another user.Do you know Xianghui Yang?You can claim authorship or link another user.Do you know Zibo Zhao?You can claim authorship or link another user.Do you know Chunchao Guo?You can claim authorship or link another user.Do you know Long Quan?You can claim authorship or link another user.

Abstract

Sparse voxel grids preserve the spatial structure needed for detailed 3D reconstruction, but their memory still grows rapidly with resolution as active surface cells increase. We introduce ChunkVAE, a sparse grid variational autoencoder organized around local chunks rather than a global latent volume. Local learned operators permit independently chosen encoder and decoder partitions and allow inference chunk sizes to differ from training. Two complementary data operators make this flexibility practical: Balanced Binary Object Partitioning distributes active cells while limiting replicated overlap, while S-Curve weighted stitching attenuates unreliable boundary features when assembling a global latent or reconstruction. Across three object benchmarks, ChunkVAE is competitive with or better than strong baselines from $512^3$ to $1536^3$; smaller chunks lower peak allocated memory and shorten per-chunk compute, enabling faster parallel inference. Stable stitched latents and improved image to 3D metrics indicate that local compression can scale geometry while retaining the global interface required downstream.

Community

00

Publication notes

Author note
14 pages, 9 figures