Harbor checks the saved answer file. EvalArc separately runs the delivered program. These three scripted controls show why the two results need distinct meanings.
Declared controls, actual container executions. Harbor 0.23.0; non-root agent; separate verifier container; eight public requests. No model inference, hidden test, or model capability ranking.
| Control | Harbor reward | Program score | All checks pass | Accepted | Evidence |
|---|---|---|---|---|---|
| reference | 100% | 100% | Yes | Yes | ATIF · Program · Answers · Import · Grade |
| clock-fault | 80% | 80% | No | No | ATIF · Program · Answers · Import · Grade |
| detached-answers | 100% | 80% | No | No | ATIF · Program · Answers · Import · Grade |
Partial credit uses the task's six published dimension weights. Version 0.2.0 introduces weighted answer scoring; the earlier 0.1.0 task retains its all-or-nothing reward. Independent acceptance here requires a valid program evaluation with score 1.0.
Summary · Preselected plan · File identities · Task contract · Harbor task · Commands and scope
The full upstream ATIF schema was checked in the recorded runtime. The offline packager checks source identity, tool linkage, recorded commands, answer scoring and evaluation consistency. It does not re-execute a program just by opening this page.
This demonstrates the stated task's answer-only verification boundary. It is not an exploit of Harbor isolation or evidence about unseen reward hacking.