A New Attempt to See Beyond Company Reports

Researchers are trying to build a more independent picture of how people actually use generative AI systems such as ChatGPT, Claude, Gemini, and Grok. A new project called the AI Observatory, co-led by Stanford Trustworthy AI Research Lab PhD candidate Anka Reuel, aggregates real AI conversations collected with user consent from seven existing datasets.
The project responds to a basic problem: the most visible accounts of AI usage often come from the companies that run the models. Anthropic and OpenAI regularly publish reports on how people use Claude and ChatGPT, but outside researchers say those reports reflect the questions and data the companies choose to share. Reuel argues that there is no independent source to corroborate them, even as policymakers and researchers make consequential judgments about AI’s benefits and risks.
What Gets Missed When Work Use Is the Focus
The AI Observatory analyzed 24,521 conversations, comprising 85,633 conversational turns, from 5,000 users interacting with 52 models between 2023 and 2025. A “conversational turn” means a user prompt and the corresponding AI response.
One of the clearest findings concerns the limits of work-focused reporting. Anthropic’s Economic Index, one of the best-known sources of AI usage data, emphasizes work and productivity uses of Claude and filters out conversations outside that scope. When AI Observatory researchers applied Anthropic’s method to their own dataset, 48% of conversations would have been excluded.
Those excluded, non-work conversations were more likely to include sensitive or personal topics:
- Health and relationships: 44.2%, compared with 31.2% in Anthropic’s analysis;
- Adult or illicit topics: 7.9%, compared with 2.1%;
- Harassment and hate: 27.5%, compared with 5.66%;
- Sexual content: 16.7%, compared with 2.4%.
OpenAI’s 2025 report similarly found that only 30% of consumer use of ChatGPT was work-related. Together, these figures suggest that productivity is only one part of the AI usage story.
Different Models, Different Behaviors

The Observatory also found that usage patterns vary across models and over time. Grok and Gemini were used more often for information retrieval. Grok was especially popular for news and politics, but misinformation also tended to concentrate there, consistent with other research; xAI did not respond to a request for comment.
Users were more likely to use Anthropic models for coding, Gemini for social and roleplay interactions, and ChatGPT for homework assistance. Even versions of the same product differed: conversations with ChatGPT powered by GPT-3.5 were shorter, while those with GPT-4o were longer and more iterative, a pattern that aligns with concerns that GPT-4o became associated with emotional dependence.
In WildChat, one of the largest and most detailed datasets in the project, conversations became longer and more elaborate over time, with increases in prompt tokens, response tokens, and conversation turns. Tokens are the small text units that AI models process. Small talk also rose, suggesting increased AI companionship, while assistants’ self-disclosure—that is, saying they are chatbots—declined. Sensitive exchanges became less frequent, which may indicate that platforms were deploying stronger safeguards.
Limited Data, Broader Access
The AI Observatory’s dataset is still tiny compared with what major AI labs can analyze internally. Anthropic’s latest Economic AI Index is based on 1 million Claude conversations, while OpenAI’s ChatGPT usage report analyzed 1.5 million conversations. Because the Observatory relies on voluntarily shared data, it may also underrepresent sensitive uses that people are less willing to disclose. The researchers caution that the findings do not represent all AI use.
Still, the project matters because it gives the broader research community another source of evidence. AI companies rarely make chat data available for independent analysis, so public reports often show only part of the picture. David Widder, an assistant professor at the University of Texas at Austin who was not involved in the project, said the Observatory’s bird’s-eye-view analysis can help researchers understand different uses more consistently, rather than leaving related findings sectioned off in separate reports.
Why This Matters for AI Governance
The study underscores that consumer AI is no longer just a productivity tool. It is also used for learning, companionship, information seeking, roleplay, coding, and sensitive personal topics. Company reports can illuminate parts of that landscape, but no single company report can capture cross-model differences or the full range of user behavior.
The likely next step is pressure for privacy-preserving data access that lets independent researchers compare systems and audit claims. Until then, decisions about AI safety, regulation, and social impact will continue to rely heavily on partial evidence. The AI Observatory is modest in scale, but it points to a larger need: independent infrastructure for understanding what people are really doing with AI.
