EVALARC / INDEPENDENT SWE REVIEW

PUBLIC SOURCE TASKS · ACTUAL GPU EXECUTION

Inspect what the grader actually established.

Follow 36 attempts across three source projects and four workflow conditions. Review the delivered skill, saved changes, tool failures and independent native verdict for each attempt.

Download the offline review and raw recordsDownload 36 structured rows

No attempt obtained upstream acceptance. Five reports have an infrastructure flag and remain uncertain. Three public tasks with repeated seeds cannot establish a general skill benefit or model ranking.

Compare delivery and outcomes

Direct and MCP workflows preload the same frozen object. Unrelated museum notes match its 444-token length and the full initial prompt length. Preloading is performed by the workflow.

ConditionAttemptsAcceptedNot acceptedUncertainNonempty patches

Each condition has three tasks × three seeds. Uncertain reports stay in the scheduled denominator. One generation request has incomplete usage.

Inspect an attempt

AttemptSource projectConditionSeedOutcomeChanged filesReview

A high pass fraction can hide a required defect.

Six native executions compare original defects with upstream fixes. The illustrative 95% aggregate rule accepts two defective controls because many regression checks still pass.

Source taskControlRequired defect testsRegression checksNative verdictAggregate ≥95%

These are scripted upstream controls, separate from the 36 model attempts. The illustrative threshold was declared after control inspection and before model generation. This comparison is not a claim of an undisclosed upstream vulnerability.

Method, uncertainty and reproduction

Task selection and generic guidance were fixed before the model cohort. Each attempt had 20 turns, up to 2,048 new tokens per turn and a 900-second interaction budget. All conditions used the same tools and immutable source images. Every scheduled slot is retained.

Repeated incorrect source paths explain many failed reads. Five pytest reports contain both offline dependency-installation failures and candidate-code errors. Their original network_unreachable label remains visible; it does not establish a single cause. The report keeps these outcomes uncertain.

The source archive includes native logs, model requests, exported patches and runtime identities. Images, weights, installed dependencies and full filesystem tar snapshots are referenced by identity, not bundled. Review records offline; reproducing execution needs the recorded runtime.

EvalArc code and upstream issue, code and test materials have separate attribution and license boundaries. See the methods and the archive's source-attribution directory.