The physical world does not care how good your model is
Digital capability curves have made the field systematically overconfident about robotics. The data constraint is structural, and it does not respond to capital.
The position
Embodiment is not behind because it is under-invested. It is behind because physical interaction data cannot be scraped, and that does not change.
Embodiment sits at 41, the lowest dimension in the Index, and it did not move this quarter. I am quoted in the counter-evidence for that dimension and I want to explain the position properly, because "robotics is behind" is usually said in a tone that implies it is catching up.
The constraint is data, and it is structural
Language models were built on a corpus humanity had already produced and left lying around. Vision models had a comparable inheritance. There is no equivalent corpus of physical interaction, and there cannot be, because the relevant data is not observations of the world — it is records of what happened when a specific body applied a specific force to a specific object.
You get that data by moving a robot. Every hour of it costs an hour, plus the hardware, plus the space, plus the breakages. Capital compresses some of that; it does not compress the hour.
41
Embodiment
What cross-embodiment transfer actually does
It is the most important result in my field in several years, and it is not a capability result. It is a data-efficiency result: data collected on one platform now serves others, which multiplies the value of every hour already spent.
That is genuinely significant. It is also not the same as the capability improving. The policies are not better; more platforms can now use them. Those are different claims and they get conflated constantly in coverage of this work, including coverage of my own.
Why nobody publishes field reliability
You have seen the demonstration videos. You have not seen a failure rate per hour in an unstructured environment, from anyone, on any platform. That absence is uniform across the industry and it is not an oversight.
The numbers are bad. Not embarrassing-bad — bad in the ordinary way that early technology is bad. But a field that publishes demonstrations and withholds reliability is a field where the outside view should assume the reliability is the reason.
Show me the failure rate per hour. Everything else is a video.
The grounding problem underneath
Multimodality scores 82 and I think that is right for perception. But the residual failures — counting, temporal ordering, precise localisation — are exactly the operations manipulation depends on. A system that describes a scene beautifully and cannot reliably say which object is nearer is not close to picking one up.
This is why I am unmoved by arguments that embodiment will be solved by the digital dimensions arriving. The digital dimensions have arrived. The specific things we need are the specific things still missing.
My actual forecast
Steady progress, no discontinuity, bounded by data collection throughput and hardware iteration cycles. Both are measured in years and neither is accelerating the way compute did. If embodiment is genuinely load-bearing for the threshold — and I think it is — then the conservative scenario deserves more weight than it currently gets.