Summary
Coverage is explicit: missing and unevaluated cells display an em dash. They are not scores.
Treatments and runner readiness
Treatment scores
The primary value uses all 74 pinned plans. Available-case values are provisional and are not leaderboard scores.
24 × 4 result matrix
Per-app results
All 74 plan results
The wedding/test1 anomaly remains in the primary universe. The separate 73-plan value is an audit diagnostic only.
Per-cell records
Prompts and upstream inputs
Every cell receives the same canonical goal prompt plus its app's pinned MVP PRD and public assets. Test plans are evaluation inputs, not construction inputs.
Prompt SHA-256:
Methodology
Process
Build agents use native subscription logins and native tools. We use the same fixed copy of the ViBench evaluator for each full score. The evaluator uses an API key. We have not yet used this evaluator to score a full ViBench test plan. The first private test can put the evaluator and the candidate app in the same container. Build agents cannot access the evaluator key. This test does not prove that the evaluator is isolated. We do not publish a comparable score until the result passes all publication checks.
Scores
Upstream points anomaly
Missing cells
Public artifacts
Data
Download the normalized public data
data.json SHA-256: 6b88f5db0090179cecbd3e8f2bddd206900596fdbc5f06a0c0a2375117d12882
This public bundle intentionally excludes candidate source, host paths, commands, raw event streams, transcripts, logs, sessions, and authentication material.