RECORDED GPU CONTEXT CONTROLS · QWEN3-8B / L40S

Check the protocol before interpreting the score.

Twelve recorded model attempts compare relevant robot-review guidance with unrelated descriptive text. Both MCP skill-load payloads contain 476 tokens, including file hashes. The original six attempts and the six-attempt follow-up remain separate.

What changed between the cohorts?

The initial programs timed out waiting for a JSON response. Their EOF example could finish without testing persistent interaction. A separate reference run passed under the same grader. The follow-up gives both conditions the same two-request protocol diagnostic and a flush instruction. It was planned after observing the original failures.

The task, model, skill texts, seeds, tools and budgets remain fixed. Initial instructions and the supplied helper differ between cohorts; outcomes are not pooled into an efficacy estimate. The helper checks JSONL interaction and does not grade numerical correctness.

01 · Initial EOF example

0 / 6 tasks resolved. 48 response timeouts in 48 recorded cases.

Each seed receives both conditions. Scores are independent program checks.
Model seedRelevant guidanceUnrelated prose
170.0%
unresolved
0.0%
unresolved
410.0%
unresolved
0.0%
unresolved
970.0%
unresolved
0.0%
unresolved

Inspect all six attempts → Download cohort

02 · Persistent-request diagnostic available

0 / 6 tasks resolved. 0 response timeouts in 48 recorded cases.

Each seed receives both conditions. Scores are independent program checks.
Model seedRelevant guidanceUnrelated prose
1787.5%
unresolved
0.0%
unresolved
4187.5%
unresolved
0.0%
unresolved
9787.5%
unresolved
0.0%
unresolved

Inspect all six attempts → Download cohort

Inspect the failure before trusting a finish signal

An agent can call finish without implementing the program. A program can return JSON correctly while computing a metric incorrectly. Each cohort preserves submitted programs, complete messages, tool receipts and independent case checks. A declared finish, program score and resolved task are separate outcomes.

The unrelated control is descriptive prose, not a second task-specific reference. It retains the same catalog name and description; inspect both exact payloads before interpreting any difference. This is one public-development task and three model seeds, not a model ranking or evidence of general skill efficacy.

Reproduce and review

Methods and exact commands · Separate cohort summaries · File identities · Reference environment check · Unchanged-program buffering diagnostic

The page displays saved records and works offline. No inference or candidate execution runs in the browser. Offline verification checks consistency; it does not authenticate a producer.