Anthropic TASTE Benchmark Tests Whether AI Can Judge Safety Research
Anthropic researchers introduced TASTE, a benchmark for judging AI safety research proposals, finding the best tested model at 60% agreement versus 77% estimated expert-human performance.
Anthropic introduces a benchmark for a hard-to-verify research skill
Anthropic researchers published TASTE — The AI Safety Taste Evaluation — on August 28, 2026 to measure whether frontier AI models can judge the quality of AI safety research proposals in ways that agree with experienced human researchers.
The work targets a difficult problem for automated AI safety research: many research decisions do not have crisp, automatically verifiable rewards. Choosing between competing research directions can depend on judgments about importance, tractability, experimental design and likely value.
TASTE therefore measures a narrow but consequential capability: whether models prefer the same empirical AI safety proposals that experienced researchers prefer. It should not be interpreted as a complete benchmark of scientific intelligence, alignment, or the ability to conduct an entire safety-research project independently.
TASTE contains 92 pairwise comparisons over 50 proposals
The paper reports 92 pairwise comparisons across 50 empirical research proposals. The researchers estimate experienced-human agreement with the benchmark labels at 77%.
To construct the proposal pool, the team began with human-written proposals from Anthropic's Fellows Program, reverse-engineered motivating questions with Claude Opus 4.6, and generated additional model-written proposals. Experienced AI safety researchers then scored and discussed proposal sets before revising their preferences.
A key benchmark-design result is that pair discussion plus filtering for self-reported strong confidence improved agreement with held-out researchers. The paper reports agreement rising from 53% before discussion to 68% for strong-confidence, post-discussion preferences, before further filtering produced the final benchmark.
Frontier models still trail the human estimate
In Anthropic's standard pairwise evaluation, Fable 5 was the best-performing tested model at 60%, compared with the benchmark's 77% estimated human-researcher performance. The paper reports that frontier-model performance spans roughly 41% to 60%.
Anthropic also reports that Fable 5 improved from 48% at low reasoning effort to 60% at maximum effort, suggesting additional test-time computation helped on this task.
However, the benchmark is small enough that model confidence intervals are wide. Anthropic notes that per-model intervals span roughly ±10 percentage points, so the results should not be used to make fine-grained claims that one nearby model is definitively better than another.
The blog also notes that almost all tested models fall within two standard deviations of chance, and that some frontier systems that perform strongly on general agentic benchmarks perform near chance on TASTE. This is evidence that research judgment may remain a distinct capability rather than automatically tracking broader benchmark strength.
The benchmark measures agreement, not objective truth
TASTE uses experienced-researcher preference as its target because many safety-research questions lack objective ground-truth answers. That makes label quality especially important.
The researchers' discussion protocol is therefore part of the contribution: evaluators first score proposals individually, then discuss disagreements in pairs, revise judgments and report confidence. The final benchmark keeps higher-confidence comparisons with larger score gaps and limits repeated use of the same proposal.
This design reduces some noise, but it does not turn subjective research judgment into an objective scientific truth. The 77% figure is itself an estimate of agreement under the benchmark's procedure, not a universal ceiling on research-quality evaluation.
Released now versus access-limited
Released now: Anthropic's Alignment Science post and the full research paper describing TASTE, its construction methodology and model evaluations.
Access-limited: the benchmark and underlying human-feedback dataset are available on request to AI safety researchers rather than as an unrestricted public download. Anthropic's page links to an access-request form.
The practical significance is that TASTE creates a concrete evaluation target for one of the hardest pieces of automating AI safety research: selecting promising research directions when there is no simple verifier to tell an agent whether its judgment was correct.
This article is built from the source material below. Open the originals for full context and the latest updates.