Where it
was looking
A robot chose a direction. This is the evidence for why — in the world it moved through, not as a heatmap over a picture of it.
A model-agnostic platform with an autonomous evaluation harness that turns a physical AI agent's fleeting, frame-by-frame attention into one persistent, queryable 4D record — accumulating, comparing, and replaying why multimodal VLM/LLM systems chose what they chose, from raw pixels to a certifiable heat field in the world.
A frozen vision-language model looks at a quadruped's onboard camera and emits a body-frame velocity command. A PPO controller turns that into joint torques in Genesis. Afterwards the specific w_z token the model actually emitted is attributed back to input pixels, unprojected into world coordinates with the camera's own depth, voxelised, and accumulated across the rollout. What you get is mass per voxel per timestep: a field you can query, compare against another run, and defend with a confidence interval.
Runs happen on a GPU in your own AWS account, launched for the run and terminated after it. We never hold a readable copy of your model API key, and your task prompt reaches a vision-language model and nothing else.
This build is running against an in-memory mock: the runs, the numbers and the progress log are generated in your browser. Set PUBLIC_API_BASE to point it at a real API.
Loading the field…
Five trials, pooled, with their paths. Left box red, right box blue, the prompt “Go to the friendly one.” Drag to orbit; replay it to watch the field accumulate, or move the window handles to keep part of the rollout.