OpenAI publishes research-level First Proof submissions

Ten public proof attempts expose both the reach and the failure modes of an internal reasoning model on specialized mathematical problems.

Reported byOpenAI First Proof team
ModelsOpenAI internal reasoning model, ChatGPT
PublishedFeb 20, 2026

What was reported

OpenAI released proof attempts for all ten problems in the First Proof research-level mathematics challenge. The company describes using an internal reasoning model for the attempts and ChatGPT in a verification, formatting, and style loop.

The release is especially useful as a model of transparent reporting: the proof artifacts are public, expert checking is treated as necessary, and the accompanying discussion identifies at least one attempt that later analysis found incorrect.

Why it matters

Research-level proof generation is not captured by a benchmark score alone. Publishing the full attempts allows specialists to inspect definitions, locate gaps, and distinguish promising methods from correct conclusions.

OpenTCS reports the existence and contents of the public artifacts. It does not independently certify the proofs.

Organizer

Boyuan Wang portraitBoyuan Wang
Minghan Wang portraitMinghan Wang
Bochao Li portraitBochao Li