PUBLIC SOURCE TASKS · ACTUAL GPU EXECUTION
Inspect what the grader actually established.
Follow 36 attempts across three source projects and four workflow conditions. Review the delivered skill, saved changes, tool failures and independent native verdict for each attempt.
This offline page contains all 36 traces. Raw records are in the adjacent attempts, controls and protocol directories one level above this review.
No attempt obtained upstream acceptance. Five reports have an infrastructure flag and remain uncertain. Three public tasks with repeated seeds cannot establish a general skill benefit or model ranking.
Compare delivery and outcomes
Direct and MCP workflows preload the same frozen object. Unrelated museum notes match its 444-token length and the full initial prompt length. Preloading is performed by the workflow.
| Condition | Attempts | Accepted | Not accepted | Uncertain | Nonempty patches |
|---|
Each condition has three tasks × three seeds. Uncertain reports stay in the scheduled denominator. One generation request has incomplete usage.
Inspect an attempt
| Attempt | Source project | Condition | Seed | Outcome | Changed files | Review |
|---|
No attempts match these filters. Choose another condition or outcome.
Saved patch
Native evaluator report
Original issue and workflow preload
Recorded usage and source identities
Model and tool interactions
The download preserves exact submitted requests, full model responses, context decisions and native logs. Displayed traces are checked against that archive.
A high pass fraction can hide a required defect.
Six native executions compare original defects with upstream fixes. The illustrative 95% aggregate rule accepts two defective controls because many regression checks still pass.
| Source task | Control | Required defect tests | Regression checks | Native verdict | Aggregate ≥95% |
|---|
These are scripted upstream controls, separate from the 36 model attempts. The illustrative threshold was declared after control inspection and before model generation. This comparison is not a claim of an undisclosed upstream vulnerability.
Method, uncertainty and reproduction
Task selection and generic guidance were fixed before the model cohort. Each attempt had 20 turns, up to 2,048 new tokens per turn and a 900-second interaction budget. All conditions used the same tools and immutable source images. Every scheduled slot is retained.
Repeated incorrect source paths explain many failed reads. Five pytest reports contain both offline dependency-installation failures and candidate-code errors. Their original network_unreachable label remains visible; it does not establish a single cause. The report keeps these outcomes uncertain.
The source archive includes native logs, model requests, exported patches and runtime identities. Images, weights, installed dependencies and full filesystem tar snapshots are referenced by identity, not bundled. Review records offline; reproducing execution needs the recorded runtime.
EvalArc code and upstream issue, code and test materials have separate attribution and license boundaries. See the methods and the archive's source-attribution directory.