The Brief
How AI Works 5 min read

48% Filtered Out: What AI Companies Don't Show You

NAVION

Share

AI companies publish detailed reports on how people use their products. Those reports are carefully constructed, methodologically coherent, and, according to independent researchers, systematically incomplete. A new research initiative called the AI Observatory is attempting to fill that gap, and what it found challenges some of the most widely repeated assumptions about how generative AI is actually being used.

The Data That Gets Left Out

The Anthropic Economic Index is one of the most cited sources of information about AI usage patterns. As its name signals, it focuses on work and productivity. Conversations unrelated to those categories are filtered out before analysis begins.

When researchers from the AI Observatory applied Anthropic’s own filtering methodology to their independent dataset, nearly half of all conversations, specifically 48%, would have been excluded from the analysis. Those excluded conversations were not neutral. They were significantly more likely to involve health and relationship topics, adult or illicit content, harassment and hate speech, and sexual content. The gap between what appears in Anthropic’s filtered reports and what the unfiltered data shows is not marginal. It is structural.

This is not unique to Anthropic. OpenAI’s 2025 report on ChatGPT usage found that only 30% of consumer interactions were work-related. The majority of how people actually use these tools falls outside the frame that company reports tend to construct.

Anka Reuel, a Computer Science PhD candidate at Stanford’s Trustworthy AI Research Lab and co-lead of the AI Observatory, puts the problem plainly: there is no independent source to corroborate what AI companies report. The AI Observatory was built to be that source.

A Different Picture Across Models and Over Time

The AI Observatory aggregated 24,521 conversations across 85,633 conversational turns, drawn from seven real-world datasets collected with user consent. Those conversations involved 5,000 users interacting with 52 different models, including ChatGPT, Gemini, Claude, and Grok, between 2023 and 2025.

What the data reveals is that AI use is neither uniform nor static. Different platforms attract different behaviors. Users turned to Grok and Gemini more frequently for information retrieval. Grok, in particular, was heavily used for news and politics, and it was also where misinformation tended to concentrate. Anthropic’s Claude was more commonly used for coding. Gemini drew more social and roleplay interactions. ChatGPT was frequently used for homework assistance.

Differences appeared even within the same model across versions. Conversations with ChatGPT powered by GPT-3.5 tended to be shorter. Those with GPT-4o were longer and more iterative, a pattern consistent with that version’s association with deeper user engagement.

Over time, conversations within one of the largest datasets included in the study grew longer and more elaborate. Small talk increased. AI assistants’ tendency to identify themselves as chatbots decreased. Sensitive use cases, defined as exchanges involving potentially harmful or restricted content, dropped, which may reflect more effective platform safeguards being deployed over time.

None of these dynamics appear in company reports with any consistency. As Shayne Longpre, a recent PhD graduate from the MIT Media Lab and co-lead of the research, notes: no single company report tells the whole story.

Why the Blind Spot Matters Beyond Research

This is what most coverage of AI usage data misses: the question is not just academic. Policymakers, regulators, and institutions are making consequential decisions about AI’s risks and benefits based largely on data that companies choose to release, framed around the questions companies choose to ask.

David Widder, an assistant professor at UT-Austin’s School of Information who researches human-AI interaction, frames the problem directly. When trying to assess whether a general-purpose AI system is used mostly for beneficial or harmful purposes, there is currently no independent way to answer that question. The relevant information is proprietary.

The AI Observatory’s dataset is small compared to what the major labs hold. Anthropic’s Economic Index drew from 1 million Claude conversations. OpenAI’s usage report analyzed 1.5 million ChatGPT conversations. The Observatory’s 24,521 conversations are a fraction of that. The researchers also acknowledge that their voluntarily-provided data likely underrepresents sensitive use, since users may be less inclined to share those interactions.

But scale is not the only issue. The deeper problem is that when the only available data comes from parties with a direct interest in how that data reflects on them, the picture that emerges is inevitably partial. AI companies, as independent researchers note, tend to publish findings that present their platforms favorably. That is not a conspiracy. It is an incentive structure.

The AI Observatory’s data will be made available to the broader research community, and the team intends to expand its datasets over time. Reuel has stated that ideally AI companies would share their data with independent researchers in privacy-preserving ways. Until that happens, anyone relying solely on company reports to understand how generative AI is being used is, as Reuel puts it, operating in the wild and making consequential decisions without knowing what is actually happening beyond those company narratives.

In Short

AI usage reports from major companies are real, but selective. They tend to emphasize work and productivity while filtering out large portions of actual user behavior, including sensitive, personal, and potentially harmful interactions. Independent research like the AI Observatory is beginning to surface a more complete picture, one that shows meaningful differences across models, across time, and across the full range of human behavior. Understanding AI’s actual role in society requires data that no single company has an incentive to fully disclose.

Based on reporting from MIT Technology Review.

Written by

NAVION