hn.today

Building Rome from a Single Image

build-rome.github.io4 points0 comments
Screenshot of Building Rome from a Single Image

A method for reconstructing a complete, metrically scaled 3D scene mesh from a single photograph, built around an object-generator redesigned to handle whole scenes. The pipeline lifts the input image into a point map and splits space into adaptive 3D chunks that grow with distance to preserve nearby detail and cover distant structures. Image features are explicitly conditioned in 2D→3D by projecting them onto observed surfaces and marking voxels as free, visible, or hidden. Each chunk is produced by a sparse-structure transformer conditioned on those lifted features and visibility, followed by a structured-latent transformer; chunks are generated autoregressively with overlapping latents reused to enforce local consistency, and a decoder assembles chunks into a unified mesh. An interactive demo exposes many reconstructions and a distance-measure tool; displayed meshes are simplified for web viewing while full-resolution outputs are used for figures and videos.

The approach argues that adaptive chunking plus explicit visibility and autoregressive latent reuse yield high-fidelity, globally consistent reconstructions from one view. Results demonstrate plausible completion of occluded geometry, preservation of fine nearby detail alongside coherent distant structure, and visual improvements over contemporary single-image scene methods in comparisons. The system produces metrically meaningful meshes useful for visualization and downstream tasks where scene-scale geometry and consistency from a single photo matter.

Read on build-rome.github.io0 comments on Hacker News

Summary generated by AI from the linked article. hn.today is not affiliated with Hacker News or Y Combinator.

More in AI

The daily digest

Today's best Hacker News stories, summarized and screenshotted, one email a day.