Usage data from the inside: what OpenAI Signals can and cannot tell us

Analysis

Usage data from the inside: what OpenAI Signals can and cannot tell us

There is no other source for this data, which is exactly why it needs reading carefully.

A magnifying glass resting on a worn paper world map

Published

July 5, 2026

Reading time

3 minutes

Perspective

Analysis

Topics

adoption · data · openai

OpenAI published data on how people around the world use ChatGPT, with country-level breakdowns of adoption and usage trends.

First-party usage data is genuinely valuable and genuinely awkward. It is the only view of this scale that exists, and it comes from a party with an interest in what it shows.

What this data is uniquely good for

Nobody else can see this. Survey-based estimates of AI adoption are unreliable — people misreport tool use, and the denominators are guesswork. Actual usage logs, aggregated properly, answer questions surveys cannot:

  • What fraction of sessions are work versus personal.
  • How usage patterns differ between countries at similar income levels.
  • Whether adoption curves flatten or keep climbing.

Those are real questions with policy implications, and this is the best instrument available for them.

The structural caveats

It is one product. ChatGPT usage is not AI usage. Countries where local providers dominate will look like non-adopters and are not. This matters most for exactly the comparisons people want to make.

Access is not uniform. Availability, pricing, payment infrastructure, and language support differ by country. A usage gap can reflect a distribution gap rather than a demand gap, and the data cannot separate them.

Category definitions are internal. When usage is bucketed into "work" and "personal," someone chose those boundaries. Reasonable people would choose differently, and the choice moves the headline numbers.

Publication is selective by default. Not through dishonesty — every organisation publishes the analyses it finds interesting. But the set of questions asked is shaped by who is asking.

How to use it well

Treat it as a high-quality measurement of one large product, not as a census of AI use. That framing keeps the value and drops the overreach.

Specifically:

  • Within-country trends over time are the most trustworthy signal, since the measurement apparatus is constant.
  • Cross-country comparisons need an availability and competition adjustment that the data cannot provide.
  • Absolute adoption rates should be read as lower bounds.

Why publishing it is still good

It would be easy for a lab to keep this entirely internal. Publishing it, with whatever framing, puts numbers into a debate that has been running largely on assertion — particularly around which countries and sectors are actually using these tools versus which are talked about.

The right response to first-party data is not to dismiss it. It is to read it with the provenance in mind, and to push for the methodology to be documented well enough that outside researchers can check the parts that are checkable.

The questions this data could answer, if opened up

There is a set of genuinely important questions that only first-party usage data can address, and that no lab has yet published on in detail:

  • Does usage substitute for or complement existing tools? Whether people use these systems instead of search, or alongside it, matters for a great deal of policy.
  • What happens to usage after the novelty period? Long-run retention curves by cohort would settle several arguments about whether adoption is durable.
  • How does usage differ where local-language support is strong versus weak? This is the clearest available evidence on whether language coverage is a real barrier to adoption.

None of these require exposing individual data, and all of them are answerable from aggregates. The reason they are not published is presumably that they are not the questions a company most wants answered publicly.

An arrangement where independent researchers could pose questions against aggregated data, under access controls, would be more valuable than any single published report — and it is the model that eventually emerged for search and social data, after considerably more friction than was necessary.

Continue reading

More from COREXA