# Meta entered its models in five STEM Olympiads. What that actually measures

- Published: 2026-08-07
- Authors: CORTEXA
- Category: Analysis
- HTML: https://researchhub-vert.vercel.app/blog/meta-models-stem-olympiads-reasoning

A perfect theory score is a real result. It is also a result on a test whose problems are written to be solvable in a fixed window by a prepared human — which is exactly what makes it easy to over-read.

Meta reported entering its models in five STEM Olympiad competitions, including a perfect score on the theory exam at the Asian Physics Olympiad.

Olympiad results are attractive because they look like the opposite of a benchmark. Nobody trains on them directly. The problems are new each year. They are graded by humans against a rubric. That is a genuinely better setup than most of what the field uses.

## Why olympiads are a good test

The useful property is not difficulty. It is **specification**. An olympiad problem states exactly what counts as a solution, and a grader who has never seen your system can tell whether you produced one. Compare that with the average agentic benchmark, where "success" is a heuristic string match against a trajectory somebody recorded once.

Olympiad problems also resist the most common form of benchmark rot. You cannot quietly inflate your score by memorising the answer key, because the answer key did not exist when your training data was collected.

## Why the result is narrower than it reads

The problems are written so that a prepared human can solve them in a fixed window — typically four to five hours, with no reference material. That constraint shapes the problems in ways that matter:

- Every problem is **known to be solvable**. Someone has verified a solution exists and fits in the time budget.
- The relevant physics is **in scope by construction**. No olympiad problem requires a result published last month.
- Problems are **self-contained**. There is no missing context to go find, no ambiguous requirement to negotiate.

Real research has none of these properties. Most of the difficulty in a research problem is establishing whether it is tractable at all, and deciding what to try when it is not.

So a perfect theory score says: on well-posed problems with guaranteed solutions, drawn from a settled body of knowledge, the model performs at the level of a strong prepared student. That is a real and non-trivial claim. It is not a claim about open-ended research.

## The part worth watching

The interesting signal in olympiad results is not the score. It is the **error profile**. A model that fails olympiad problems the way a strong student fails them — running out of time, misreading a setup, making an algebra slip — is doing something different from a model that fails by confidently producing a physically impossible answer.

That distinction does not show up in a headline number, which is why the score alone is the least informative part of the announcement.

## What to ask of any olympiad claim

1. **Theory or experimental?** These are separate exams testing different things. A perfect theory score is not a perfect overall score.
2. **How many attempts?** Best-of-n on a graded exam is a different result from single-shot.
3. **What tooling was available?** A model with a calculator and a scratchpad is not the same system as one without.
4. **Was the grading independent?** Self-graded olympiad results are a much weaker claim.

None of that diminishes the achievement. It just locates it. Olympiad performance is evidence about a specific, well-defined capability, and the field has a persistent habit of reading such evidence as being about something much larger.
