iPhone, 360 Camera or LiDAR: What Your Capture Guarantees, Not What It Renders

by Pierre
iPhone, 360 Camera or LiDAR: What Your Capture Guarantees, Not What It Renders

For eighteen months everyone has been comparing renders: put two splats side by side, see which one looks better. That is the wrong question. The difference is not in the final quality — it is in what the capture guarantees before you leave the site.

Two different physics

Two weeks ago, writing about Atlas by World Labs, we looked at AI learning to generate Gaussian splats. This piece is about the other half of the problem: capturing them from the real world. If the format itself is new to you, we covered what a Gaussian splat is separately.

On a handsome façade, on an overcast day, with a careful capture, an iPhone holds its own against a $5,000 scanner. So the disagreement is not about the result. It is about the process.

The iPhone does photogrammetry. You film, you extract frames, and a Structure-from-Motion algorithm — COLMAP, GLOMAP, RealityScan — reconstructs the camera position at each instant after the fact, by matching points across views. 3DGS training only begins once those poses are estimated. The whole chain rests on one assumption: there is enough texture, enough overlap and enough photometric stability for matching to converge.

LiDAR measures. A beam sweeps the scene, an inertial unit tracks the motion, a camera array supplies colour, and SLAM fuses it all in real time. Poses are not inferred from pixels: they come out of the trajectory. Geometry is not guessed, it is acquired, at metric scale, while you walk.

The distinction sounds academic. It decides everything that follows.

The pipeline, end to end

Recent iPhone360 cameraXGRIDS PortalCam
Hardwarealready in your pocket€400–600$4,999
On site10–30 min, lock everythingone slow walk, single passcontinuous walk, single pass
Camera posesSfM after the factSfM after the fact, far safer convergencereal-time SLAM, on board
Scalearbitrary, realign by handarbitrarynatively metric
Recovery after failurego back to siterarepractically never needed

The table says the essential thing. With the iPhone, the uncertainty sits after the capture: you come home with a video and you do not yet know whether it will produce anything. With the scanner, it is resolved during, and you watch it happen. Turnaround follows: half a day to two days on one side, a few hours on the other.

What the iPhone does better

Do not write it off. It wins on several fronts, and not minor ones.

Close-up detail. A 48 MP sensor fifty centimetres from an object resolves a fineness of texture no 12 MP fisheye can match. For a product, a sculpture, a façade element shot up close, it remains superior.

Object scale. Orbiting a chair, a vase, a pack: photogrammetry is at home there. A LiDAR calibrated for ranges of several dozen metres adds almost nothing.

Zero marginal cost. Reshooting because the light turned costs nothing. No battery to charge, no kit to carry, no licence to activate: you are already equipped, permanently.

What the scanner does better

Interiors. This is the sharpest tipping point. White walls, corridors, uniform surfaces, mirrors, glazing: all configurations where SfM loses correspondence and returns wrong poses, or no poses at all. LiDAR does not care — it measures a distance, not a resemblance.

Indoor/outdoor continuity. On an iPhone you have to treat the two separately: the exposure gap and the texture break derail global alignment. SLAM walks straight through.

Large volumes. The PortalCam captures 856,000 points per second out to 60 metres. Roughly 280 m² across several floors in a fifteen-minute walk. An equivalent iPhone capture needs hundreds of images, splitting into blocks, then registering the blocks together.

Repeatability. Models come out metric, with relative accuracy around two centimetres. Above all: the same scene captured twice gives a usable result twice. For a client deliverable, that predictability is worth more than a few percent of sharpness.

The third way: the 360 camera

This is the option most comparisons forget, and probably the best effort-to-result ratio today. A consumer 360 camera applies exactly the PortalCam’s optical principle — spherical coverage, continuous walk — without LiDAR or SLAM, for a tenth of the price.

What it changes. Every frame sees the whole scene. Overlap between views becomes enormous, loops close naturally, and SfM converges where a conventional capture drops out — typically indoors. The three orbits at different heights disappear: one slow walk is enough. Time on site collapses, almost to scanner levels.

What it costs. Angular resolution, and that is the only real trade-off. An 8K equirectangular spreads its pixels over 360°×180°, about 22 pixels per degree, against roughly 55 for a 4K iPhone across a 70° field. A factor of 2 to 2.5 in texture fineness. On a room volume, painless. On a detail at one metre, visible.

The pipeline. Do not push raw equirectangular into COLMAP: extract perspective views or cubemap faces, declared as a coherent rig, then SfM, then training. LichtFeld Studio, open source, reads Insta360 .insv containers directly. Note that Postshot, long the free reference, has moved to a subscription: its free tier no longer exports .ply. And since August, the Insta360 X6 ($699) removes the pipeline entirely — the splat comes out of the app.

The obstacles, because there are some

Let us stay clear-eyed. Each route has its blind spot, and none is the one advertised.

On the iPhone side. No loop closure: over a long path, drift accumulates uncorrected. Moving vegetation poisons matching, glass generates floating artefacts, and forgotten auto-exposure is enough to ruin an entire sequence. Most failures are not algorithm failures, they are discipline failures.

On the scanner side. The PortalCam aims at visual fidelity, not survey work: its point cloud is not usable in topography. Processing goes through LCC Studio, which still wants a decent NVIDIA workstation — “no PC” only covers acquisition — and the LCC format remains proprietary.

On the 360 side. The tripod and the operator, which have to be masked. Parallax artefacts in the seam between the two fisheyes, awkward at close range. Small sensors that fall apart quickly in dim interiors — the format’s most underrated weakness. And for the X6’s spatial capture, a detail that matters depending on your clients: processing goes to a third-party vendor in Shenzhen, with no public schedule, and several markets do not have access yet.

The decision grid

Three questions usually suffice.

Object or space? An object, a detail, a single well-textured room: iPhone. A walkable space, a sequence of rooms: 360. A whole building, indoor/outdoor continuity: scanner.

Internal or client deliverable? To explore and feed a demo, the iPhone is unbeatable. As soon as there is a delivery date and a single window of site access, the question closes: you cannot afford to discover the next day that SfM did not converge.

How many times a year? This is the calculation that settles it. One capture a quarter: 360 covers the need for the price of a business lunch. One a month with commercial stakes behind it: the reshoot time saved pays off the scanner within a year.

And a fourth, which is not exclusive: why choose? Capture structure and interiors in 360, go back with the iPhone on detail zones, both sets feeding the same reconstruction. You get the coverage of one and the fineness of the other. What you still will not have, and only LiDAR gives: metric scale, and certainty before leaving the site.

The essential point: a guarantee beats sharpness

The iPhone is not the scanner’s rough draft, and the scanner is not a better iPhone. Three tools, three distinct constraints: fineness at low cost, coverage at low cost, guaranteed results at scale.

For two years the question has been:

“Which tool produces the best-looking splat?”

As these captures become dated deliverables rather than demonstrations, the question becomes:

“What does my capture guarantee before I leave the site?”

The wrong reflex is to choose by looking at screenshots. The right one is to start from the deliverable — what the scene has to make possible, and what happens if the capture fails.


Sources: XGRIDS PortalCam · Insta360 X6 Spatial Capture · LichtFeld Studio · COLMAP

Related Content