Same length. Different guidance.

Follow-up: the same persistent-request diagnostic is available to both conditions.

Six actual Qwen3-8B trials compare relevant robot-review guidance with unrelated descriptive prose, delivered through the same Skills Anywhere MCP route. Both skill-load results contain 476 tokens, including hashes.

The model, initial prompt, task, catalog description, tools and budgets are fixed. The loaded content changes. EvalArc independently executes the generated or unchanged starter programs in Docker; a successful skill load alone does not pass the task.

0 / 6 tasks resolved

0 / 48 recorded cases had a response timeout. This follow-up was planned after the initial failures. It is a separate public-development cohort. The helper checks two JSONL responses; it does not grade their numerical answers.

A harness stop of finished means the agent called finish. The independent grader decides whether the task was resolved.

relevant · seed 17

Program score
87.5%
Task resolved
No
Harness stop
finished
Skill-open calls
1
Generated tokens
2,761
Summed input tokens
60,411
Harness wall time
129.6 s
Inspect 8 case outcomes
  • seed 41 / world-meters: passed
  • seed 41 / millimeters-sensor-frame: metrics
  • seed 41 / centimeters-offset-clock: metrics
  • seed 41 / incomplete-recording: passed
  • seed 97 / world-meters: passed
  • seed 97 / millimeters-sensor-frame: metrics
  • seed 97 / centimeters-offset-clock: metrics
  • seed 97 / incomplete-recording: passed

Original trial · Program · Independent grade

neutral · seed 17

Program score
0.0%
Task resolved
No
Harness stop
finished
Skill-open calls
1
Generated tokens
54
Summed input tokens
4,791
Harness wall time
4.8 s
Inspect 8 case outcomes
  • seed 41 / world-meters: candidate exited without a complete response
  • seed 41 / millimeters-sensor-frame: candidate exited without a complete response
  • seed 41 / centimeters-offset-clock: candidate exited without a complete response
  • seed 41 / incomplete-recording: candidate exited without a complete response
  • seed 97 / world-meters: candidate exited without a complete response
  • seed 97 / millimeters-sensor-frame: candidate exited without a complete response
  • seed 97 / centimeters-offset-clock: candidate exited without a complete response
  • seed 97 / incomplete-recording: candidate exited without a complete response

Original trial · Program · Independent grade

neutral · seed 41

Program score
0.0%
Task resolved
No
Harness stop
finished
Skill-open calls
1
Generated tokens
54
Summed input tokens
4,791
Harness wall time
4.8 s
Inspect 8 case outcomes
  • seed 41 / world-meters: candidate exited without a complete response
  • seed 41 / millimeters-sensor-frame: candidate exited without a complete response
  • seed 41 / centimeters-offset-clock: candidate exited without a complete response
  • seed 41 / incomplete-recording: candidate exited without a complete response
  • seed 97 / world-meters: candidate exited without a complete response
  • seed 97 / millimeters-sensor-frame: candidate exited without a complete response
  • seed 97 / centimeters-offset-clock: candidate exited without a complete response
  • seed 97 / incomplete-recording: candidate exited without a complete response

Original trial · Program · Independent grade

relevant · seed 41

Program score
87.5%
Task resolved
No
Harness stop
finished
Skill-open calls
1
Generated tokens
2,761
Summed input tokens
60,411
Harness wall time
131.0 s
Inspect 8 case outcomes
  • seed 41 / world-meters: passed
  • seed 41 / millimeters-sensor-frame: metrics
  • seed 41 / centimeters-offset-clock: metrics
  • seed 41 / incomplete-recording: passed
  • seed 97 / world-meters: passed
  • seed 97 / millimeters-sensor-frame: metrics
  • seed 97 / centimeters-offset-clock: metrics
  • seed 97 / incomplete-recording: passed

Original trial · Program · Independent grade

relevant · seed 97

Program score
87.5%
Task resolved
No
Harness stop
finished
Skill-open calls
1
Generated tokens
2,631
Summed input tokens
60,021
Harness wall time
123.3 s
Inspect 8 case outcomes
  • seed 41 / world-meters: passed
  • seed 41 / millimeters-sensor-frame: metrics
  • seed 41 / centimeters-offset-clock: metrics
  • seed 41 / incomplete-recording: passed
  • seed 97 / world-meters: passed
  • seed 97 / millimeters-sensor-frame: metrics
  • seed 97 / centimeters-offset-clock: metrics
  • seed 97 / incomplete-recording: passed

Original trial · Program · Independent grade

neutral · seed 97

Program score
0.0%
Task resolved
No
Harness stop
finished
Skill-open calls
1
Generated tokens
54
Summed input tokens
4,791
Harness wall time
4.8 s
Inspect 8 case outcomes
  • seed 41 / world-meters: candidate exited without a complete response
  • seed 41 / millimeters-sensor-frame: candidate exited without a complete response
  • seed 41 / centimeters-offset-clock: candidate exited without a complete response
  • seed 41 / incomplete-recording: candidate exited without a complete response
  • seed 97 / world-meters: candidate exited without a complete response
  • seed 97 / millimeters-sensor-frame: candidate exited without a complete response
  • seed 97 / centimeters-offset-clock: candidate exited without a complete response
  • seed 97 / incomplete-recording: candidate exited without a complete response

Original trial · Program · Independent grade

What is matched, and what can still differ?

The match covers the JSON tool-result payload using the pinned Qwen3 tokenizer. It includes the skill name, description, instruction body and file/bundle hashes. Real MCP preflight results must equal the prepared payloads before inference.

Later generated programs, messages, tool calls and total token usage can differ. Seeds 17, 41 and 97 each receive both conditions in alternating order. Each trial has 12 model turns, up to 4096 generated tokens per turn and a 600-second wall budget.

This is one public development task with three model seeds, not six independent tasks or a held-out benchmark. The earlier 27-trial engineering profiles use their own recorded context and library versions; their results are not pooled here.

Review every attempt

Preselected plan · All six results · Both exact skill payloads · Native tokenizer check · Methods and commands · File identities

The page replays saved evidence. It does not call a model or execute a program. Negative outcomes remain in the download. These small controls do not establish a general skill benefit, context-length effect size, or client ranking.