Real model calls / NVIDIA L40S / Public development evidence
A skill loaded.
Did the task pass?
Trace a robot recording from numerical observations to a graded answer. Compare no skill, direct file loading and real MCP delivery, with every attempt available to inspect.
Loading a skill is observable. Following it is a separate question. These small engineering pilots show no skill accuracy gain; changed prompts and discovery rules prevent pooling profiles into one benchmark.
Loading recorded evidence…
01 / Choose an engineering profile
Select a cell to inspect its receipt. Percentages show independent partial score; full resolution requires every check to pass.
02 / Follow one recorded attempt
Actual tool sequence
Model, skill pins & file identities
Read the result in context
Discovery improved. Accuracy did not.
The final profile's six direct/MCP trials all open the skill and finish their workflow. None passes independent grading. Two of three no-skill trials fully resolve the task. These are three seeds on one public task, not a ranking of models, agents or MCP.
The first two profiles are also retained, including search loops, exhausted turn budgets and incomplete candidates. Candidate code runs in non-root containers without network or host mounts. The model's finish message never determines acceptance.