Tuesday, August 25, 2026

Independent technology reporting and practical analysis

Manila ·
ARTIFICIAL INTELLIGENCE

Independent reporting, useful context, and practical analysis.

Back to Technomalist
Artificial Intelligence / news

Independent AI Observatory Exposes Gaps in Vendor Usage Reports

A new public dataset of real AI conversations suggests company reports miss significant personal and sensitive uses, from health to harassment.

Featured image for Independent AI Observatory Exposes Gaps in Vendor Usage Reports
Featured image for Independent AI Observatory Exposes Gaps in Vendor Usage Reports

MIT Technology Review reported on a new independent project called the AI Observatory that aims to fill a gap in how AI companies disclose usage data. Companies like Anthropic and OpenAI regularly publish reports on how people use Claude and ChatGPT, but researchers say those reports only show what the companies choose to release. Anka Reuel, a PhD candidate at Stanford's Trustworthy AI Research Lab, said there is no independent source to corroborate the data.

The AI Observatory, co-led by Reuel and MIT Media Lab PhD graduate Shayne Longpre, aggregates and analyzes real AI conversations from seven existing datasets, collected with user consent. It includes 24,521 conversations involving 5,000 users and 52 different models, including ChatGPT, Gemini, Claude, and Grok, spanning 2023 to 2025. The platform is intended to give researchers and policymakers a more independent basis for assessing how people actually use generative AI.

When the team applied Anthropic's Economic Index methodology—which filters for work- and productivity-related conversations—to their data, they found that 48% of conversations would have been excluded. Those filtered-out conversations were more likely to involve health and relationships (44.2% versus 31.2% in Anthropic's analysis), adult or illicit topics (7.9% versus 2.1%), harassment and hate (27.5% versus 5.66%), and sexual content (16.7% versus 2.4%). OpenAI's 2025 report similarly found that only 30% of consumer use was work-related.

Anthropic has separately published blog posts about companionship and even CSAM generation, but David Widder, an assistant professor at UT-Austin's School of Information who was not involved with the Observatory, said having a bird's-eye view across all uses in one dataset helps researchers understand the differences more consistently.

The researchers also found that AI use changed over time. Conversations in the WildChat dataset became longer and more elaborate, with more small talk and less self-disclosure from assistants that they were chatbots—suggesting increased companionship use. Sensitive use cases, including sexual harassment and hate speech, dropped over time, which may indicate better safeguards.

Usage varied significantly by model. People used Grok and Gemini more frequently for information retrieval; Grok was especially popular for news and politics, but it was also where misinformation tended to concentrate. This is consistent with other research, and xAI did not respond to a request for comment. Users turned to Anthropic's Claude for coding, Gemini for social and roleplay uses, and ChatGPT for homework assistance. Even versions of the same model differed: conversations with GPT-3.5 were shorter, while GPT-4o conversations were longer and more iterative—a version later associated with emotional addiction.

The Observatory's dataset is small compared to what the large labs hold. Anthropic's latest Economic Index analyzed 1 million conversations, and OpenAI's report analyzed 1.5 million. Anthropic said its published research reflects its teams' specific questions and that it supports external independent research. OpenAI did not respond to requests for comment.

Because the data comes from voluntarily provided sources, sensitive uses are likely undercounted, and the researchers caution that the findings are not indicative of all AI use. Still, the project broadens access for the research community, since AI companies do not typically share raw chat data. Widder said that without such access, we cannot answer whether a system is used mostly for good or bad. Reuel added that anyone making decisions based on AI usage data risks operating "completely in the wild" without knowing what is actually happening beyond company narratives. The Observatory's data will be available to researchers, and the team hopes to expand it over time.

See an error? Read our corrections policy or email [email protected].

MORE FROM TECHNOMALIST

Continue reading

View all
Featured image for Researchers Show Expired Visa Contactless Cards Can Be Revived for Fraud
Cybersecurity

Researchers Show Expired Visa Contactless Cards Can Be Revived for Fraud

Philips LatteGo 4400 Series espresso machine on a kitchen counter.
Guides

Philips LatteGo 4400 espresso machine drops to AU$613 on Amazon Australia

Alice talks to Nora and Frank
Entertainment

How AI and assistive tools are helping disabled actors like Steve Way thrive on 'Furious'