← PlaygroundRobot Reel × Skills Anywhere × EvalArcMethods & limits

Real model calls / NVIDIA L40S / Public development evidence

A skill loaded.
Did the task pass?

Trace a robot recording from numerical observations to a graded answer. Compare no skill, direct file loading and real MCP delivery, with every attempt available to inspect.

Loading a skill is observable. Following it is a separate question. These small engineering pilots show no skill accuracy gain; changed prompts and discovery rules prevent pooling profiles into one benchmark.

27complete trials
3 × 3conditions × seeds per profile
ATIF 1.8upstream schema checked

Loading recorded evidence…

Read the result in context

Discovery improved. Accuracy did not.

The final profile's six direct/MCP trials all open the skill and finish their workflow. None passes independent grading. Two of three no-skill trials fully resolve the task. These are three seeds on one public task, not a ranking of models, agents or MCP.

The first two profiles are also retained, including search loops, exhausted turn budgets and incomplete candidates. Candidate code runs in non-root containers without network or host mounts. The model's finish message never determines acceptance.