Skip to content

PreviewAll content, scores and forecasts here are illustrative sample data — not reporting, and not measurements.What this means

Agitology
Perspective8 min readSample

The physical world does not care how good your model is

Digital capability curves have made the field systematically overconfident about robotics. The data constraint is structural, and it does not respond to capital.

The position

Embodiment is not behind because it is under-invested. It is behind because physical interaction data cannot be scraped, and that does not change.

Dr. Soma VargaRoboticist · Independent

Embodiment sits at 41, the lowest dimension in the Index, and it did not move this quarter. I am quoted in the counter-evidence for that dimension and I want to explain the position properly, because "robotics is behind" is usually said in a tone that implies it is catching up.

The constraint is data, and it is structural

Language models were built on a corpus humanity had already produced and left lying around. Vision models had a comparable inheritance. There is no equivalent corpus of physical interaction, and there cannot be, because the relevant data is not observations of the world — it is records of what happened when a specific body applied a specific force to a specific object.

You get that data by moving a robot. Every hour of it costs an hour, plus the hardware, plus the space, plus the breakages. Capital compresses some of that; it does not compress the hour.

41

Embodiment

Flat across two quarters while every other dimension in the Index rose.

What cross-embodiment transfer actually does

It is the most important result in my field in several years, and it is not a capability result. It is a data-efficiency result: data collected on one platform now serves others, which multiplies the value of every hour already spent.

That is genuinely significant. It is also not the same as the capability improving. The policies are not better; more platforms can now use them. Those are different claims and they get conflated constantly in coverage of this work, including coverage of my own.

Why nobody publishes field reliability

You have seen the demonstration videos. You have not seen a failure rate per hour in an unstructured environment, from anyone, on any platform. That absence is uniform across the industry and it is not an oversight.

The numbers are bad. Not embarrassing-bad — bad in the ordinary way that early technology is bad. But a field that publishes demonstrations and withholds reliability is a field where the outside view should assume the reliability is the reason.

Show me the failure rate per hour. Everything else is a video.

The grounding problem underneath

Multimodality scores 82 and I think that is right for perception. But the residual failures — counting, temporal ordering, precise localisation — are exactly the operations manipulation depends on. A system that describes a scene beautifully and cannot reliably say which object is nearer is not close to picking one up.

This is why I am unmoved by arguments that embodiment will be solved by the digital dimensions arriving. The digital dimensions have arrived. The specific things we need are the specific things still missing.

My actual forecast

Steady progress, no discontinuity, bounded by data collection throughput and hardware iteration cycles. Both are measured in years and neither is accelerating the way compute did. If embodiment is genuinely load-bearing for the threshold — and I think it is — then the conservative scenario deserves more weight than it currently gets.

Continue

Other perspectives