Initial cohort: an EOF example is the supplied execution check.
Six actual Qwen3-8B trials compare relevant robot-review guidance with unrelated
descriptive prose, delivered through the same Skills Anywhere MCP route.
Both skill-load results contain 476 tokens, including hashes.
The model, initial prompt, task, catalog description, tools and budgets
are fixed. The loaded content changes. EvalArc independently executes the generated
or unchanged starter programs in Docker; a successful skill load alone does not pass the task.
0 / 6 tasks resolved
48 / 48 recorded cases had a response timeout. An EOF example can appear to work while a persistent JSONL service waits on buffered output. The original failures remain here; the subsequent cohort supplies a protocol diagnostic to both conditions.
A harness stop of finished means the agent called finish. The independent grader decides whether the task was resolved.
The match covers the JSON tool-result payload using the pinned Qwen3 tokenizer.
It includes the skill name, description, instruction body and file/bundle hashes.
Real MCP preflight results must equal the prepared payloads before inference.
Later generated programs, messages, tool calls and total token usage can differ.
Seeds 17, 41 and 97 each receive both conditions in alternating order. Each trial
has 12 model turns, up to 4096 generated tokens per turn and a 600-second wall budget.
This is one public development task with three model seeds, not six independent
tasks or a held-out benchmark. The earlier 27-trial engineering profiles use their
own recorded context and library versions; their results are not pooled here.
The page replays saved evidence. It does not call a model or execute a program.
Negative outcomes remain in the download. These small controls do not establish
a general skill benefit, context-length effect size, or client ranking.