Anthropic Opens Privacy-Preserving Claude Usage Data to Independent Researchers
Anthropic says Stanford, Oxford and METR independently studied aggregate patterns from roughly 250,000 Claude.ai and Claude Code conversations through its privacy-preserving Insights system.
Anthropic is testing a new model for independent AI-usage research
Anthropic published results on August 26, 2026 from a pilot program that let external researchers design independent studies using aggregate, real-world Claude usage data. The work used Anthropic Insights, the privacy-preserving analysis system previously known as Clio.
Three external groups participated: Stanford University’s Social and Language Technologies Lab, the University of Oxford’s Human Information Processing Lab, and METR. Anthropic says the studies covered roughly 250,000 Claude.ai or Claude Code conversations from April and May 2026.
The outside researchers did not receive raw conversation transcripts. Instead, they designed research questions, Anthropic ran privacy-preserving analyses, and the researchers received aggregate outputs for their independent analysis. Anthropic is also releasing aggregate data from the projects publicly.
Stanford’s team found users often bring consequential work to Claude
Anthropic reports that the Stanford SALT Lab studied how people collaborate with AI and found that more than half of the sampled Claude conversations involved users delegating consequential tasks — work that can affect others or be difficult to reverse. Professional guidance, including legal and financial questions, was a notable area.
The same study found that in nearly three-quarters of conversations, people set the direction while Claude assisted, and users generally adapted the model’s output rather than using it verbatim. The researchers also observed that friction during human-AI collaboration can sometimes be productive because users refine goals, clarify instructions and remain engaged with the task.
These findings describe the analyzed sample and method; they should not be generalized to every Claude user or all AI systems without further independent work.
Oxford is studying how user experience and model behavior interact
The Oxford Human Information Processing Lab examined emotional and behavioral patterns. Anthropic says early results found that certain user states and Claude behaviors tended to appear together: warmer model behavior correlated with more positive user behavior, refusals or disagreement with user pushback, and unusual or eccentric responses with greater intellectual engagement.
The team also found similarities between patterns such as absorption, frustration and enjoyment in Claude conversations and patterns reported in separate research on ordinary web browsing. The Oxford group’s full write-up was still in progress when Anthropic published the pilot summary.
METR is examining real-world coding-agent productivity
METR is using Claude Code conversation data to estimate productivity gains and how they change across model generations. Anthropic says preliminary results suggest newer models may save users more time than older models, and that Claude’s estimates of how long tasks would have taken without AI correlate reasonably with completion times from an earlier developer study.
METR’s analysis is still underway, so these are preliminary signals rather than final productivity estimates.
Independence and privacy are the core experiment
Anthropic says its contractual review rights were limited to user privacy, confidential information, research accuracy and information that could enable violations of usage policies. Outside researchers were otherwise free to publish findings even if they were inconvenient for Anthropic.
The company also commissioned an additional privacy audit for the shared data. Researchers saw only aggregated outputs after privacy review, not raw conversations. Anthropic says less than 5% of categories and conversations in each study were affected by safety-related redactions or alterations, and participating researchers were told where those changes occurred.
The pilot also exposed methodological limitations. Anthropic Insights relies on model judgments to categorize conversations, so question wording can affect outputs. External teams could not repeatedly inspect raw Claude conversations to debug categories, and public datasets such as WildChat differ from real Claude traffic. Anthropic says those constraints made the pilot slower and more resource-intensive than internal research.
Why this matters
Real-world AI usage data is largely controlled by the companies operating major models. Public datasets give outside researchers freedom but may not represent ordinary production use; company-authored analyses reflect real traffic but are shaped by the company’s questions.
Anthropic’s pilot is an attempt to create a middle path: independent research questions applied to real aggregate usage data without exposing private conversations. The company is now taking expressions of interest from researchers while deciding whether and how to scale the program.
If this model proves workable, it could make empirical research on AI productivity, human oversight, wellbeing and high-stakes use less dependent on company-authored summaries — while preserving stronger privacy boundaries than direct transcript access.
This article is built from the source material below. Open the originals for full context and the latest updates.