2026-07-20

What Autonomous Embodied Intelligence Needs

Anthropic just drove real robots with an off-the-shelf coding agent [1]. The lesson that reorganizes the field: the same model swings from near-zero to working when you change the control interface, not the weights. Read carelessly, that says the model is fine and only the scaffolding is missing — and that reading is half right and dangerously tidy. Autonomous embodied intelligence (AEI) has two bottlenecks at once: a reasoner whose multi-step, multimodal grounding is genuinely good enough, and a set of parts around it that barely exists for the physical world. Here is the shape of both — figures first.

01The shape

L  — safety · authority · approval · audit  (cross-cutting) world obs C · context base model multi-turn · multimodal E · loop T · tools body model-facing world-facing S · state read / write persistent memory — across episodes geometry · semantics · robot state · task progress · skills V  — verification & trajectory record: is it working?  (cross-cutting)

Not a bigger model, not a better policy — a reasoner wired to a body through six parts.

Whatever you call the layer around the model — the agent-systems world calls it a harness [3]; we care about the parts, not the label — AEI needs six: an execution loop (E), a tool-and-policy registry (T), context management (C), cross-episode state (S), lifecycle safety hooks (L), and verification (V). A coding agent gets all six free from a fifty-year-old text workspace: the filesystem is its state, tests are its verifier, a stack trace localizes failure, git revert undoes anything [4]. A robot inherits none of it. Its world owns a clock it cannot pause, hides its state behind noisy partial observation, and never throws a stack trace — so three of the six (E's clock, C's attention, V's ground truth) do not just need rebuilding, they get qualitatively harder [5].

02We tried it — and only half of it works

✔  WORKS

  • fable 5, zero-shot VLN — no navigation training
  • Just a simple waypoint action space + multimodal dialogue
  • A general reasoner navigates out of the box

✘  NOT DONE

  • Habitat: constant wall-sliding
  • VLNVerse: ~10% collision rate
  • A frontier model on every step isn't affordable

This month we let a coding agent driving fable 5 attempt vision-and-language navigation zero-shot — no VLN training, just a waypoint action space and a multimodal dialogue. It works surprisingly well: a general reasoner, wrapped in a thin loop with the right interface, is a competent navigator out of the box. That is the existence proof — the parts, not a bespoke model, carry the task. But we will not be drunk on it. The demo runs; AEI does not — yet. The six parts are the distance between those two sentences.

EExecution loop — two clocks, hybrid now

Embodiment runs on two clocks [7]. The rollout loop ticks on the environment's clock — one step per tick, whether or not anything reasons; the agent loop (ReAct [8]) ticks on the reasoner's — one LLM turn per tick, no environment required. Today's navigators sit at the two broken extremes: NavGPT-style workflows [9] script every step (total but rigid), and our zero-shot agent is pure ReAct (flexible, but its unhandled corners are the wall-sliding). Anthropic sharpens it: real-time control wants ~83 Hz while inference runs at 0.2–0.4 Hz, and more reasoning per step buys almost nothing. The deployable answer is hybrid — a fast structured skeleton with the reasoner invoked on events (stuck, uncertain, interrupted); the horizon past it is full-duplex, controller and reasoner free-running on their own clocks. Build the hybrid so it can grow into the full-duplex one.

structured — the world waits concurrent — the world never waits workflow NavGPT — total, rigid pure ReAct ours — partial, wall-slides hybrid skeleton + event-triggered reasoner full-duplex both clocks free-run, coupled by buffers deployable today = the middle; horizon = the right edge

TTools & policies — interface beats model

The registry decides what the model can actually do, and it dominates. Direct control fails (LIBERO 0–5.5%; no model stands a humanoid up); hand the same model a pretrained policy and it works. We see the slope from the bottom too: a waypoint predictor lifts a 4B model from 0.00 to 0.30 navigation success — no new weights, just a better interface. The one caveat is selection, not capability: cut a tool set from 15 to 2 and accuracy went 80%→100% [3] — a hundred tools is a hundred-page manual nobody reads fast. So prepare abundantly, expose selectively. And to be clear: more and better policies do not compete with the loop — they are what it steps and calls. Growing the policy zoo and building the loop are the same project.

direct torque / pose · ≈0–5.5% programmatic write a controller RL supervision design reward, train a policy policy control command a pretrained policy · ≫ rising abstraction & success →    + waypoint predictor: 4B nav 0.00 → 0.30

CContext — attention, not capacity

Context management is not "fit more in the window" — it is attention. Anthropic's cleanest result: depth maps, segmentation, and crosshairs are useless or harmful, while a compass (heading angle as text) and a cursor (a queryable 3D point) lift every model — one climbs 6%→32% on a 10-task subset. The bottleneck was never "can't see"; it was "doesn't know where it is pointed." So C has two jobs: before the window fills, place the right representation where attention lands; after it fills, compress. A million-token window retires neither — long contexts dilute attention, the "lost in the middle" effect [13].

helps — act-on-able state compass → heading angle (text) cursor → queryable 3D point the model can locate itself in it doesn't help — raw pixels depth map · segmentation crosshair overlay nothing to orient against compass + cursor: one frontier model 6% → 32%

SState — the part that breaks first at home

Everything above lives in one episode; state is what survives across them, and it breaks first at household scale. A navigation episode is hundreds of steps; "keep the kitchen tidy this week" is millions, with no reset — no window holds that, so even fable 5 chokes. Competence has to live outside the transcript, in durable state: a persistent map, an object memory, task progress, a skill library. In-context learning will not carry it — it decays within the window and flips sign with scale (our demo episode: a 4B model improved 0→1/10, but a 27B model got worse, 3/10→0/8). The implication is blunt: do not leave durable competence in the context; write it to state, and eventually to weights. S is also the deepest attack surface — a poisoned map persists across every session until writes are validated. ICL flips sign · 4B 0→1/10 · 27B 3/10→0/8

LLifecycle hooks — authority the world can't undo

In code, lifecycle hooks are a convenience; in a robot they are execution semantics, because the world has no undo. Force and speed limits, collision and proximity gates, workspace bounds, approval for irreversible acts — these decide whether a one-way door opens. Anthropic's real-robot crashes are exactly hook failures: one model judged a trash can "safe, it is left of the crosshair," drove in, and dragged it two meters; another charged a glass door, mistaking its reflection for the target. Both are perception errors — and both a proximity-and-speed gate would have caught without consulting the model at all. That is the point: L is the one part that holds even when the reasoner is confidently wrong, and it is where the human belongs — uncertainty past a threshold should stop and ask, not act. both crashes = gate failures

VVerification — success without a stack trace

A coding agent knows when it succeeded — green tests, a stack trace naming a line. A robot knows almost nothing of the kind: success is latent, delayed, uncertain, and failure could be perception, pose, planner, controller, or skill choice. Wall-sliding is the canonical unverified failure — the agent believes it is progressing because nothing contradicts it. So V must manufacture the signal the world withholds: progress monitors, geometric and physics checks, world-model or simulator rollouts that test a plan before committing, human feedback as backstop. And V does double duty: a good verifier turns each rollout into a labeled trajectory — the fuel for everything next. wall-slide = the unverified failure

06Where it goes — the flywheel

Put the six together and a direction falls out. The large model is the right call now — only a strong multimodal reasoner carries a task zero-shot. But paying for it on every step is not necessary, and its own limits say it is not the endpoint: the destination is small, post-trained policies, with the big model kept for the hard, novel, supervisory moments — which is what Anthropic's VLA results actually show, the frontier model's value being to judge when the policy would fail, not to out-act it. This is where the six parts become an engine: good parts let you run the big model over many policies (T), keep it oriented and bounded (C, E), hold state (S), stay safe (L), and — through V — label every trajectory it produces. Those labels post-train the next small policy, which runs cheaper and collects still better data. That is a flywheel, and good parts plus post-training spin it.

good components E · T · C · S · L · V run big model over many policies, in the world V labels trajectories verified → data S persists them durable, recallable post-train small policy cheaper · specialized data flywheel
Autonomous embodied intelligence is not a bigger model or a better policy. It is a capable multimodal reasoner coupled to a body through six parts — loop, tools/policies, context, state, safety, verification — none of them free from the world. Two bottlenecks, not one: fix the reasoner and the parts. Hybrid now, small post-trained policies next, full-duplex on the horizon.
References

[1] Anthropic, "Claude plays robotics (Embody)," 2026. anthropic.com/research/claude-plays-robotics.

[2] Allen Institute for AI, "MolmoAct," 2025. allenai.org/blog/molmoact.

[3] J. Meng et al., "Agent Harness for Large Language Model Agents: A Survey," Preprints, 2026.

[4] J. Zhou, "Inside the Coding-Agent Harness," 2026. companion post.

[5] J. Zhou, "What It Takes to Build a ReAct-Style Embodied Agent," 2026. companion post.

[7] J. Zhou, "Rollout Loop vs. Agent Loop," 2026. companion post.

[8] S. Yao et al., "ReAct," ICLR, 2023. arXiv:2210.03629.

[9] G. Zhou, Y. Hong, Q. Wu, "NavGPT," AAAI, 2024. arXiv:2305.16986.

[13] N. F. Liu et al., "Lost in the middle," TACL, 2024. arXiv:2307.03172.

← Back to blog