AutoR&D-Engineer, tested on AutoResearchExam benchmark

We plug GPT-6 Astra into our AutoResearch framework and compare it with published Astra runs using Terminus 2 on Bespoke Labs’ AutoResearchExam benchmark.

We follow Bespoke Labs’ scoring protocol: original task environments and graders, no internet access in task containers, and validation-based checkpoint selection with separate hidden-test scoring. We employ an adversarial reward-hacking judge to review task environments and completed runs.

Choose a task to explore

Make every decoding step do less work.

Generate text faster on an eight-core CPU while meeting the benchmark’s token-agreement requirement.

Task brief

Watch the research unfold

Best validation reward · higher is better

AutoR&D-Engineer · AstraTerminus 2 · Astra
Recorded checkpoint progressCompare the selected local research campaign with the published Astra run. Use the research-time slider or Replay research button to explore the trajectory.

Loading recorded checkpoints…

24h 00m
Recorded checkpoints; scores are held between evaluations.

Four more tasks. More ideas to explore.

Final results within ±5% of published Astra. Open a task to see the ideas, checkpoints, and evidence.

Click a card to explore

Read the result. Explore the evidence.

Recorded checkpoints, the original task graders, and validation-based selection. Each result links the research decisions to the measurements behind them.

What does the progress curve show?

Validation progress tracks the best validation reward reached so far. Hidden-test progress tracks the hidden reward of that same validation-selected solution. The curves use the benchmark website’s difficulty-adjusted reward mapping, so higher is better on every chart. They are drawn as steps from recorded evaluations, with no invented intermediate improvements.

Final performance and AUARC

CPU decoding, decoder graphs, sparse embeddings, and subset selection highlight the final hidden-test metric. Categorical learning and covariance estimation highlight 24-hour hidden-test AUARC: the average reward held over time. Finding a good solution earlier contributes more area. The percentages compare the metric named on each result card.

Model and reasoning settings

Both systems use GPT-6 Astra. The six campaigns shown here use high reasoning effort; Bespoke Labs reports using the highest effort available for its evaluations.

Evaluation methodology and safeguards

We use the benchmark’s original task graders and difficulty-adjusted reward mapping. Checkpoints are selected using validation scores, with separate hidden-test grading of the selected solutions. Retained artifacts and result files are linked by checked hashes.

We employ an adversarial reward-hacking judge to review task environments and completed runs. Our records also include source, task-constraint, and access reviews.

Read the benchmark methodology
Task coverage and comparison sources

We randomly sampled CPU tasks to keep Astra evaluation costs manageable. The initial campaign covered 21 CPU tasks; seven H100-only tasks were excluded. This page highlights six selected task stories and four additional near-ties from those evaluations.

Near-ties use the relative difference in each task’s final hidden-test metric, with a ±5% threshold. The comparison data was captured on September 21, 2026.

At this checkpoint

Validation determines selection. Hidden-test measurements report how that selected solution generalizes.

Read the original task

Ideas that moved the run forward

Open an idea to see what changed and how it measured.