Anthropic Funds Independent AI Wellbeing Evaluations
Anthropic launched a $5 million program to fund independent, open-source evaluations of how AI systems affect user wellbeing across realistic multi-turn conversations.
Anthropic is funding independent wellbeing evaluation research
Anthropic announced on August 25, 2026 a $5 million grant program for independent research into how AI systems affect users' wellbeing. Selected grantees are expected to build open-source evaluations that can be used across the AI industry, with Anthropic providing direct funding, model access and technical support.
This is an evaluation and research-funding initiative, not a new Claude model or a claim that current wellbeing safeguards are solved. Anthropic says funded researchers will work independently and publish their work as open-source projects.
The program targets long-horizon conversational risk
Anthropic argues that wellbeing is difficult to measure from a single model response because risk can emerge over a longer interaction. A response that is appropriate in isolation can become inappropriate when earlier messages reveal escalating distress, dependency, disordered eating or another sensitive context.
The company is therefore asking for evaluations that represent realistic multi-turn conversations, including scenarios where risk changes over time. This is an important distinction from static safety benchmarks that score one prompt-response pair without conversation history.
Anthropic sets methodological expectations for funded evaluations
Its published guidance asks evaluation developers to define clearly what is being measured and why a pass or failure matters; involve clinicians and other subject-matter experts in design and validation; test both insufficient safeguards and excessive refusal; reflect real user behavior; and validate automated graders against human experts.
Testing both harms and precautions matters because a wellbeing system can fail in two directions: it can respond too freely in a risky situation, or become so restrictive that it refuses reasonable, beneficial conversations.
Anthropic says the goal is to expand the field beyond the company's internal safeguards and invite psychologists, clinicians, methodologists and other specialists to develop stronger measurement standards. Applications for the program are due September 21, 2026, with selected applicants invited to submit full proposals.
What the announcement does and does not establish
The grant program does not itself demonstrate that an evaluation reliably predicts real-world wellbeing outcomes. Benchmark quality will depend on scenario design, clinical validity, population coverage, cultural context, grader agreement and whether benchmark performance transfers to deployed conversations.
It also does not replace crisis-response policies, product safeguards, human support pathways or external clinical research. The program's value will depend on the rigor and practical adoption of the open evaluations that grantees ultimately release.
For AI safety teams, the initiative is notable because it moves wellbeing evaluation toward independent, expert-validated and longitudinal testing instead of relying only on short, internally designed safety prompts.
This article is built from the source material below. Open the originals for full context and the latest updates.