We follow Bespoke Labs’ scoring protocol: original task environments and graders, no internet access in task containers, and validation-based checkpoint selection with separate hidden-test scoring. We employ an adversarial reward-hacking judge to review task environments and completed runs.
Choose a task to explore
Make every decoding step do less work.
Generate text faster on an eight-core CPU while meeting the benchmark’s token-agreement requirement.
Every improvement has a story. Watch an idea appear at each new validation best.
Loading recorded checkpoints…
Recorded checkpoints; scores are held between evaluations.
Four more tasks. More ideas to explore.
Final results within ±5% of published Astra. Open a task to see the ideas, checkpoints, and evidence.
Click a card to explore ↘
Read the result. Explore the evidence.
Recorded checkpoints, the original task graders, and validation-based selection. Each result links the research decisions to the measurements behind them.
What does the progress curve show?+
Validation progress tracks the best validation reward reached so far. Hidden-test progress tracks the hidden reward of that same validation-selected solution. The curves use the benchmark website’s difficulty-adjusted reward mapping, so higher is better on every chart. They are drawn as steps from recorded evaluations, with no invented intermediate improvements.
Final performance and AUARC+
CPU decoding, decoder graphs, sparse embeddings, and subset selection highlight the final hidden-test metric. Categorical learning and covariance estimation highlight 24-hour hidden-test AUARC: the average reward held over time. Finding a good solution earlier contributes more area. The percentages compare the metric named on each result card.
Model and reasoning settings+
Both systems use GPT-6 Astra. The six campaigns shown here use high reasoning effort; Bespoke Labs reports using the highest effort available for its evaluations.
Evaluation methodology and safeguards+
We use the benchmark’s original task graders and difficulty-adjusted reward mapping. Checkpoints are selected using validation scores, with separate hidden-test grading of the selected solutions. Retained artifacts and result files are linked by checked hashes.
We employ an adversarial reward-hacking judge to review task environments and completed runs. Our records also include source, task-constraint, and access reviews.
We randomly sampled CPU tasks to keep Astra evaluation costs manageable. The initial campaign covered 21 CPU tasks; seven H100-only tasks were excluded. This page highlights six selected task stories and four additional near-ties from those evaluations.
Near-ties use the relative difference in each task’s final hidden-test metric, with a ±5% threshold. The comparison data was captured on September 21, 2026.