3D Geodata Academy

VGGT drift fix, splats in Rerun, Emesent | 3D Intelligence Report, August 18, 2026

Published on August 18, 2026 • 41 min read

← All 3D Intelligence Reports

3D Intelligence Report – August 18, 2026

Theme of the day: the reconstruction stack and the navigation stack are turning into one stack. The paper fixes feed-forward 3D reconstruction so a long sequence stops drifting, and it reports the win in absolute trajectory error, which is a robotics metric, not a rendering one. A humanoid company pays up to $400K for classic visual-inertial SLAM. An underground mapping firm gets funded to turn GPS-denied navigation into a platform other people's robots can run. A logging viewer puts Gaussian splats on the same synced timeline as lidar sweeps and camera poses. And the venue where most of this work lands first opened its call.

Every link below was fetched and verified on August 18, 2026, the day this report went out.

Paper published 2026-08-15

VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction

accepted to ACM Multimedia 2026 (Rio de Janeiro, November 10 to 14) · arxiv 2608.15260

Wei Zhang, Yihang Wu, Songhua Li, Qi Wang

up to 32% lower absolute trajectory error than chunk-based alignment on long sequences

VGGT has become the default feed-forward 3D backbone this year, but it only sees a short window at a time. Run it on a long video and the chunks have to be stitched, which normally means solving a 7 degree-of-freedom Sim(3) fit per join: rotation, translation, and scale. Scale is the weak leg, so the fit is ill-conditioned and the error compounds down the sequence. VGGT-Align pulls scale-defining geometric invariants (ground plane and structural cues) straight out of the predicted point clouds, which pins the scale before alignment and collapses the join to a well-conditioned 6 degree-of-freedom rigid transform. A light test-time self-supervised refinement pass cleans up what is left. Evaluated on KITTI, Waymo, OpenLORIS and virtual KITTI.

Florent’s take

What I like here is that they went after the conditioning of the problem instead of throwing more network at it. Everyone building on VGGT hits the same wall: it is sharp inside its window and it forgets across windows, and the stitching is where your reconstruction quietly bends. Fitting scale, rotation and translation together at every join means the shakiest of the three drags the other two around. Find something in the scene that already tells you the scale, and you are solving a much better behaved problem. Notice the metric they report too. Not PSNR, not chamfer distance. Absolute trajectory error. This is a reconstruction paper that has decided it is a localization paper.

Job verified open 2026-08-18

Helix AI Engineer, Localization and Mapping

Figure AI · San Jose, California (onsite)

$200,000-$400,000 base (disclosed)

Florent’s take

Read the requirements rather than the headline. Feature matching, Structure from Motion, nonlinear optimization, sensor fusion. That is a photogrammetry and geodesy skill set with a humanoid robotics job title on it. The interesting part is that a humanoid has a harder version of the problem than a car does: the sensors are bolted to a body that walks, so the camera rig is shaking, the joint encoders and the IMU disagree, and there is no wheel odometry to fall back on. Figure is not hiring generalist robotics people for this, they carved out localization and mapping as its own discipline and priced it accordingly. If you have spent years fighting drift on a mobile scanner, this is the same fight with better funding.

Opportunity deadline checked on the call page

IEEE ICRA 2027 call for technical papers

Global · deadline Sep 15, 2026 · 28 days left

no funding, call for papers

Florent’s take

ICRA is not where a geospatial person usually thinks of publishing, and that is exactly why I am putting it here. A large share of the 3D perception, mapping and scene understanding work that ends up in our pipelines lands at ICRA and IROS first, months before it reaches the remote sensing journals. Twenty-eight days is a real runway for work you already have sitting in a drawer, eight pages, double-anonymous, no grant paperwork. The cost of trying is a paper. The catch worth knowing up front: acceptance means flying to Seoul in May, because they require in-person presentation, so budget for that before you submit.

Tool released 2026-08-10

Rerun 0.36.0

Gaussian splats become a first-class archetype next to point clouds and poses

Rerun is the open-source multimodal logging and visualization SDK a lot of robotics and 3D vision people already use to watch their pipelines run. 0.36.0 adds an experimental GaussianSplats3D archetype with a PLY loader, plus an EWA (elliptical weighted average) splatting path with spherical harmonics support inside re_renderer. So a splat scene out of COLMAP, gsplat or nerfstudio now loads into the same time-synced viewer that is already holding your point clouds, camera poses and sensor streams, instead of needing its own dedicated splat viewer in another window.

Florent’s take

The feature that matters is not the renderer, it is the shared timeline. Splats have lived in their own viewers since they arrived, which means comparing a splat reconstruction against the trajectory that produced it has meant two windows, two coordinate conventions and a lot of squinting. Putting the splat on the same clock as the lidar sweep and the camera pose turns a rendering artifact back into a measurement you can interrogate. Experimental means experimental, so do not build a deliverable on it this month. But this is the direction I have wanted for two years.

Market announced 2026-08-12

A $2M CRC-P grant to build Cortex AI, an open hardware-agnostic autonomy platform for GPS-denied navigation, mapping and SLAM-based robot operation

Emesent

The grant comes from Australia's CRC-P program, with Emesent leading alongside Queensland University of Technology and EPE. Cortex AI takes the SLAM and autonomy stack proven in Emesent's Hovermap and GX1 products and repackages it as a hardware-agnostic layer for aerial and ground robots plus vehicle automation, aimed at mining, construction, defence and infrastructure.

Florent’s take

Emesent has spent years selling a payload. This grant is about selling the part inside it. GPS-denied navigation, which is to say mines, tunnels, plant rooms and building interiors, is the hardest layer sitting under most digital twin work, and until now the way you got it was to buy somebody's box. If the autonomy stack becomes a layer you run on your own hardware, the scanner drops another notch toward commodity and the value moves to what you do with the trajectory and the map. That is the same story as the Topcon and GreenValley move I covered on July 28, from the other end of the stack. Worth deciding early which side of that line your business sits on.

Get the next report in your inbox

Five verified finds, my take on each, one short email a day.

Ready to start?

2,800+ engineers started exactly here. Free. No credit card. No tricks.

Get Instant access

✓ Used by teams at Meta, Airbus, CNRS

Architect Spatial Intelligence.

The Brain-to-Deploy methodology. From first principles to production-grade 3D AI.

The Foundation
€1 997 €7 249
Founding Price Lifetime Access

Master the core 3D AI stack. For innovators building a strategic edge in spatial intelligence.

  • Spatial Accelerator (17 Episodes) The 17-episode deep-dive on the Brain-to-Deploy methodology. From mental models to shipped production systems.
  • Full 3D Course Library (20+ Courses) The complete curriculum. Point clouds, meshing, segmentation, deep learning 3D, spatial reasoning.
  • Neurones 3D Software Suite The software suite I built to unify 3D reconstruction, segmentation, and spatial analysis in one stack. Standard commercial license included.
  • Monthly Spatial AI Nuggets A monthly briefing on what moved in 3D AI. Research, code, and market signals I think you should know.
  • Private Job & Market Intelligence
Secure Foundation Access
Best Value
Professional
€2 997 €13 997
Founding Price Lifetime Production Access

Scale from prototype to industrial-grade deployment. For founders architecting proprietary spatial systems.

  • Everything in Foundation
  • 4 OS Deep-Dive Production Tracks Complete tracks: Spatial Reconstructor, Segmentor, Deep Learning 3D, and more. Each ships with 12 months of active updates and support, with optional annual renewal. Each track valued at €1,497.
  • 5 Forge .exe Apps + AI Agent Toolkit Five ready-to-run Windows apps plus the AI agent toolkit. Run 3D AI pipelines without touching infrastructure code.
  • 12-Month Strategic Production Briefs Every month I spot a real market opportunity and hand you the full step-by-step blueprint to build it. Twelve briefs, twelve shipped tools or software over the year.
  • Monthly Live Q&A Sessions Monthly live calls. Ask anything technical or strategic and get a direct answer from me.
Accelerate to Professional
Architect
€4 997 €19 249
Strictly Limited: 15 Seats Direct Access

Elite 1-on-1 advisory. I personally review your architecture and deliver custom Brain-to-Deploy blueprints.

  • Everything in Professional
  • Onboarding + Annual Strategy Sessions A kickoff session to map your system, plus recurring strategy calls where I pressure-test your architecture and roadmap.
  • Private 48h Priority Channel A direct line to me for architecture reviews and technical unblocks. Guaranteed response within 48h on business days.
  • Co-Built Custom 3D AI Solution Not built for you, but with you. I code alongside you to architect and ship a custom 3D AI solution for your specific use case.
  • Portfolio & Project Endorsement
Apply for Architect Advisory

Reach out or book a call
to keep learning

Reach out for tailored support, or book a call to have more information about new courses.

Scroll to Top