OpenAI and the APA: what an evidence body actually adds
Labs can measure what their systems do. They are poorly positioned to establish what a healthy outcome looks like for a fifteen-year-old.

Published
July 15, 2026
Reading time
2 minutes
Perspective
Policy
Topics
safety · policy · openai
OpenAI announced work with the American Psychological Association on evidence-based guidance and safeguards for youth mental health and AI use.
Partnership announcements are easy to dismiss. This one is worth taking seriously for a structural reason: it involves a body whose institutional incentives are not aligned with the product.
The gap a lab cannot close alone
A model developer can measure a great deal — what the system outputs, how often it refuses, how users respond in the session, what people report afterwards. That is genuinely useful data and no one else has it.
What a developer cannot supply is the outcome standard. Whether a given interaction pattern is healthy for an adolescent over months is a question with an existing research literature, existing methods, and existing disagreements. It is not answerable from product telemetry, and a lab that tried to answer it alone would be marking its own homework on the one question that matters most.
Professional bodies have the opposite profile. They have the longitudinal research tradition and the clinical grounding, and no revenue attached to the answer. They also, typically, have no access to the systems in question.
Why the combination is the point
The two failure modes are symmetric:
- Lab acting alone: excellent measurement, self-serving standards.
- Outside body acting alone: credible standards, no visibility into the system being judged.
Combining them is the obvious move. It is also harder than it sounds, because the useful version requires the outside body to have real access and real ability to publish findings the lab dislikes.
What would make this substantive
The announcement is an announcement. The things that would distinguish substance from positioning are concrete and checkable:
- Does the APA get access to real interaction data, under appropriate protections, or only to summaries the lab prepares?
- Can they publish independently, including findings unfavourable to the product?
- Does the guidance produce specific product constraints — age handling, escalation paths, refusal behaviour — or general advice?
- Is there a mechanism for the guidance to change the product after launch, not just before it?
None of these are visible yet, and it is reasonable to withhold judgement until they are.
The reason to pay attention anyway
Youth mental health is one of the few areas where the field is likely to face binding regulation, and where the regulation will be written by people who are not machine learning researchers. Work that establishes shared evidence standards early tends to shape what that regulation looks like.
That is true regardless of the motives behind any particular partnership — which is the practical argument for watching what this produces rather than debating why it was announced.
Continue reading