Follow-up: the same persistent-request diagnostic is available to both conditions.
Six actual Qwen3-8B trials compare relevant robot-review guidance with unrelated
descriptive prose, delivered through the same Skills Anywhere MCP route.
Both skill-load results contain 476 tokens, including hashes.
The model, initial prompt, task, catalog description, tools and budgets
are fixed. The loaded content changes. EvalArc independently executes the generated
or unchanged starter programs in Docker; a successful skill load alone does not pass the task.
0 / 6 tasks resolved
0 / 48 recorded cases had a response timeout. This follow-up was planned after the initial failures. It is a separate public-development cohort. The helper checks two JSONL responses; it does not grade their numerical answers.
A harness stop of finished means the agent called finish. The independent grader decides whether the task was resolved.
The match covers the JSON tool-result payload using the pinned Qwen3 tokenizer.
It includes the skill name, description, instruction body and file/bundle hashes.
Real MCP preflight results must equal the prepared payloads before inference.
Later generated programs, messages, tool calls and total token usage can differ.
Seeds 17, 41 and 97 each receive both conditions in alternating order. Each trial
has 12 model turns, up to 4096 generated tokens per turn and a 600-second wall budget.
This is one public development task with three model seeds, not six independent
tasks or a held-out benchmark. The earlier 27-trial engineering profiles use their
own recorded context and library versions; their results are not pooled here.
The page replays saved evidence. It does not call a model or execute a program.
Negative outcomes remain in the download. These small controls do not establish
a general skill benefit, context-length effect size, or client ranking.