51 Researchers Behind Preprint Claiming Physics Benchmarks Are Miscalibrated
A new preprint on arXiv with 51 listed authors claims that leading physics benchmarks for frontier models are "broken" as evaluations and near saturation — but so far, the claim exists only in the article's title.

51 Researchers Behind Preprint Claiming Physics Benchmarks Are Miscalibrated
A new preprint on arXiv with 51 listed authors claims that leading physics benchmarks for frontier models are "broken" as evaluations and near saturation — but so far, the claim exists only in the article's title. AIMag has no confirmed details beyond the bibliographic facts, and no peer-review status is confirmed in the available documentation.
What Is Actually Verified
The verified facts are narrowly bibliographic, and they come from the arXiv listing itself (arXiv:2609.13009). The article exists on arXiv, it was submitted on September 11, 2026, the registration number is arXiv:2609.13009, and the author list is given as Ali Ansari and 50 other authors — 51 people in total.
Everything else is the authors' own claim as it appears in the title. According to the listing's title, the authors argue two things at once: that expert-based re-grading shows leading physics evaluations miscalibrate frontier models, and that these evaluations are near saturation, meaning they can barely distinguish between leading systems anymore. AIMag has verified that the title advances this claim, but has no confirmed information about the study's method, which benchmarks were re-graded, which models were tested, or any numerical results. No details beyond the title and metadata level are available in the material provided for this article.
A Title-Level Claim, Not an Established Finding
How the main claim is framed matters. "Broken evaluations and near saturation" is not a published, peer-reviewed finding at this point; it is a claim the authors advance in the article's title. The preprint carries no confirmation of peer review in the available documentation, and AIMag is not aware of any independent replication of the results. The substance of the article itself — abstract, method, and findings — has not been reviewed by AIMag, and the claim therefore can neither be confirmed nor refuted here.
If the claim holds, the consequences would be significant. Benchmark numbers carry real weight in the AI industry: labs use them in product communication, investors use them to assess progress, and policymakers use them to reason about capabilities and risk. A documented failure of leading physics benchmarks to distinguish frontier models — or a documented pattern of grading errors — would call a widely used measurement tool into question, and the same question could be raised for other fields with similar evaluation setups.
But this chain of reasoning currently rests on a title. How the authors' expert-based re-grading was designed, what criteria it used, and how large any grading errors actually are cannot be assessed from the material available here.
Why the Article Still Matters Now
The preprint lands in the middle of an ongoing debate about what frontier models' top scores on demanding evaluations actually mean. That a group of 51 researchers led by Ali Ansari claims, in their title, that leading physics benchmarks are "broken" as evaluations and near saturation is in itself a remarkable contribution to that debate — regardless of the article's ultimate fate. It signals that at least one large research group considers the common reading of physics benchmark tables unreliable.
The Open Questions
Several things must happen before the claim can be treated as established. First, the full article — abstract, method, which benchmarks and models are involved, and the numerical findings — must be reviewed; none of those details appear in the documentation provided here. Second, the peer-review status should be confirmed when and if it changes. Third, independent commentary, or responses from the labs whose models were assessed, would help test whether others can replicate the re-grading's approach and results.
For now, the correct reading is that a large group of researchers has advanced a title-level claim that leading physics benchmarks miscalibrate frontier models and can barely distinguish between them. It is a documented contribution to the evaluation debate — and, at this stage, nothing more than that.