RECORDED GPU CONTEXT CONTROLS · QWEN3-8B / L40S
Twelve recorded model attempts compare relevant robot-review guidance with unrelated descriptive text. Both MCP skill-load payloads contain 476 tokens, including file hashes. The original six attempts and the six-attempt follow-up remain separate.
The initial programs timed out waiting for a JSON response. Their EOF example could finish without testing persistent interaction. A separate reference run passed under the same grader. The follow-up gives both conditions the same two-request protocol diagnostic and a flush instruction. It was planned after observing the original failures.
The task, model, skill texts, seeds, tools and budgets remain fixed. Initial instructions and the supplied helper differ between cohorts; outcomes are not pooled into an efficacy estimate. The helper checks JSONL interaction and does not grade numerical correctness.
0 / 6 tasks resolved. 48 response timeouts in 48 recorded cases.
| Model seed | Relevant guidance | Unrelated prose |
|---|---|---|
| 17 | 0.0% unresolved | 0.0% unresolved |
| 41 | 0.0% unresolved | 0.0% unresolved |
| 97 | 0.0% unresolved | 0.0% unresolved |
0 / 6 tasks resolved. 0 response timeouts in 48 recorded cases.
| Model seed | Relevant guidance | Unrelated prose |
|---|---|---|
| 17 | 87.5% unresolved | 0.0% unresolved |
| 41 | 87.5% unresolved | 0.0% unresolved |
| 97 | 87.5% unresolved | 0.0% unresolved |
An agent can call finish without implementing the program. A program can
return JSON correctly while computing a metric incorrectly. Each cohort preserves submitted
programs, complete messages, tool receipts and independent case checks. A declared finish,
program score and resolved task are separate outcomes.
The unrelated control is descriptive prose, not a second task-specific reference. It retains the same catalog name and description; inspect both exact payloads before interpreting any difference. This is one public-development task and three model seeds, not a model ranking or evidence of general skill efficacy.
Methods and exact commands · Separate cohort summaries · File identities · Reference environment check · Unchanged-program buffering diagnostic
The page displays saved records and works offline. No inference or candidate execution runs in the browser. Offline verification checks consistency; it does not authenticate a producer.