3D Intelligence Report – September 29, 2026
Theme of the day: keeping track of what you can't see. A 31-author paper trains a video world model to remember objects that slide out of view and releases 1.5 million training clips to do it. Bedrock Robotics hires an engineer to keep autonomous excavators sure of their own position on ground they are digging away. CloudCompare 2.14 early beta stops asking for a GPS time shift on LAS files and adds three plugins. AMD agrees to buy World Labs for 8.2 billion dollars in stock.
Every link below was fetched and verified on September 29, 2026, the day this report went out.
Training Object Permanence in World Models
Haotian Zhang, Fengyuan Yu, Dezhi Luo and 28 co-authors including Philip Torr, Alan Yuille, Lvmin Zhang, Yilun Du and Hokin Deng
1.5M clips across 150 occlusion tasks
Video generators are the most common kind of world model today, and they routinely forget an object once it's occluded: a ball rolls behind a box and never comes out, or comes out as something else. The authors turn a classic developmental-psychology test into a data factory. 150 hand-designed tasks in six cognitive categories are rendered in Blender with randomised speed, lighting and camera, 10,000+ samples per task, while each task keeps its underlying structure. They release a 1.5M-sample training corpus and a 300-question exam, evaluate 14 video models (3 reference-to-video, 7 edit, 4 continuation) in a blind pairwise Elo study, and train PWM-WROP, a 16B model built on NVIDIA Cosmos3-Nano, which ranks first among continuation models and third overall. 295 upvotes on Hugging Face papers this week.
What I like here is that they didn't invent a new architecture, they wrote down what a model should know and generated the evidence at scale. That's the move for anyone in 3D: if your model fails a case, make the case reproducible in a renderer and turn it into data. For point cloud and scan work, occlusion isn't a corner case, it's most of the scene, so a world model that forgets what's behind a column is useless on a site. Two cautions. The licence is non-commercial, and synthetic Blender scenes are a long way from a cluttered plant room.
Software Engineer, State Estimation & Localization
USD 170,000 to 230,000 base (est.)
One line in this posting should stop every surveyor: the machines estimate their position in environments they are actively reshaping. On a normal job site the ground is your reference. Here the bucket removes the reference every thirty seconds. That's a harder problem than road autonomy, and it's the one construction actually has. And the skills list reads like a geodesy syllabus with C++ attached: GNSS, IMU, calibration, time sync, honest error estimates. The pay isn't published, so my range is an estimate from comparable roles this month.
CloudCompare v2.14 early beta
3 new plugins, no more GPS time shift
Published as a pre-release, 'v2.14 early beta', on 2026-09-28 at 11:16 UTC. Scalar fields natively handle large values, so no GPS time shift is needed when loading LAS files, and names can exceed 256 characters. Display speed of clouds and meshes improved through a composite GLSL program, a normals LUT texture and a colour-scale texture. New plugins: G3Point (granulometry, from the Rennes lidar team), VoxFall (non-parametric volumetric change detection for rockfalls, with CSV report) and 3D Forest Inventory (moved from Python to a standalone C++ plugin). New dotBIM import, colour filters (Gaussian, bilateral, median, mean), circle and disc entities. The command line gains -MATCH_SCALES, -FACETS (qFacets headless, shapefile and CSV export), -DISTANCES_FROM_SENSOR, -SCATTERING_ANGLES, -STAT_FIT and RGB/SF filters. LAS saving now maps non-standard scalar fields to extra bytes and lets you choose the LAS offset; E57 keeps sensor information and timestamps on a round trip. Point-pair alignment moves to the Umeyama algorithm. Oculus support is dropped in favour of a side-by-side stereo mode.
The line I'd underline is the boring one: LAS GPS time loads without a shift. Anyone who has watched timestamps round to garbage in a float field knows how many workflows quietly broke on that. The command-line additions matter more than the GUI ones for me, because -FACETS and -MATCH_SCALES headless means CloudCompare slots into a scripted pipeline instead of an afternoon of clicking. VoxFall is the plugin I'd try first on monitoring data. It's an early beta, so keep 2.13 for anything you deliver this month.
Agrees to acquire World Labs, Fei-Fei Li's spatial-intelligence lab, for about 8.2 billion dollars in an all-stock transaction
AMD's release, dated 2026-09-28 16:05 EDT, frames the deal as model expertise for chip design: the World Labs team brings research that helps AMD build hardware, software and systems for reasoning, robotics, simulation and physical AI workloads. Fei-Fei Li becomes Executive Vice President and Chief Scientist reporting to Lisa Su, and the World Labs team keeps its focus on model research. Close is expected by the end of 2026, subject to regulatory approvals. TechCrunch notes the two companies already had an inference and training partnership. It is the largest price yet put on a spatial-intelligence lab, and it pulls world models into the GPU rivalry with NVIDIA, which already ships open world models (Cosmos).
I read this as a chipmaker buying a map of the next workload. You can't design silicon for world models from benchmarks alone, you need people who know where the model actually spends its time, and that's what 8.2 billion dollars of stock buys. For us, the question is practical: what happens to Marble and to the export paths people built on it. The release says nothing on that, so I'd keep your own data and formats portable until it does. And notice today's paper trained a 16B world model on Amazon's Trainium chips; the hardware lock-in era for this kind of work may be shorter than people think.
Get the next report in your inbox
Five verified finds, my take on each, one short email a day.
Five verified finds with my take, one short email a day.