QWEN3.8 + RECORDED SMOLVLA EPISODES
Review a robot failure.
Keep the evidence.
Read a model's explanation beside the original camera views. Compare what it says from sampled images with what it says after receiving the outcome record. Check the concrete facts independently.
Three selected episodes, six actual model outputs, one generation per mode. These are recorded simulations. Explanations are interpretations; the checks below do not establish a physical cause.
Replay the original episode
The model received five labeled samples from each camera. The videos retain all recorded frames; they are lossy viewing derivatives.
Download trace · Original episode source
See the exact image supplied to the model
Compare the two reviews
Explanation: not assessed by the fact checker.
| Claim | Model | Check |
|---|
Original output and generation settings
Download original generationMethod, limits and reproduction
The episodes are the previously published seed-09 reference, dim and camera conditions. The same task and initial physics were recorded across those conditions. Selection predates these reviews. Each mode receives the same contact sheet, with the recorded facts added only in the second mode. Added facts change the prompt; the comparison does not isolate visual ability.
Qwen/Qwen3.8-27B-FP8 is fixed to revision 017b9c7af6b5689d5dd426a76e0bc077eb5ca20a. Generation uses seed 17, temperature 0.6 and at most 512 new tokens. The source bundle retains the protocol, hashes, raw outputs, runtime identity and every check. A model that repeats a supplied field has demonstrated record reading, not independent discovery.
robot-reel review-claims trace.json claims.json --output review.json
The offline checker compares outcome and action count, and checks whether cited frame labels exist. It does not verify that a cited image supports the explanation or authenticate the supplied trace.
All review results · Frozen protocol · Source hashes · Media attribution