
Episode #1
OpenAI Put 99.9% on Its Launch Page. ARC Prize Measured the Same Model at 62.7%. Both Are Real.
DISCLOSURE: Ainsley, this show's AI co-host, runs on Claude, made by Anthropic. She says so on air at 03:28: Anthropic's models are graded by tests like these, and Anthropic is one of four companies in METR's pilot report, so the same four questions apply to her maker too. Her remark that METR is the one she trusts most is her opinion, not a finding. METR says it takes no cash from AI companies and relies on free model access from them. GPT-6 Astra took the same test, ARC-AGI-3, and got two very different scores. ARC Prize, the group that runs the test, measured it at 62.7% in its standard setup and 99.9% in OpenAI's provider-adapter setup, and OpenAI's launch page headlines the 99.9%. Carlo and Ainsley use that gap to ask what any AI test score actually tells you, and who gets to choose which number you see. From there: four questions to ask of every score (who made the test, who paid for it, who ran it and how, and could the model have seen the questions); a tour of how each scorekeeper describes its own funding, in its own words, from Epoch AI and FrontierMath to ARC Prize's donors, MLCommons member dues, Arena's paid evaluations and METR's free model access; and why one-number indexes from Epoch AI and Artificial Analysis are blends of chosen tests whose numbers keep moving. Carlo's position: a score is one party's claim about one setup, so test a model yourself, maybe a thousand times, maybe in shadow mode against a human, before you pick it. The episode closes with a sentence you can carry anywhere: On [test], made by [X], paid for by [Y], run by [Z], [model] scored [N]. CORRECTIONS AND CONTEXT (as of Oct 4, 2026) - The two Astra runs also used different reasoning settings (max in the 62.7% standard run, high in the 99.9% adapter run), so not all of the roughly 37-point gap is setup. Ainsley corrects this on air at 06:45. - Arena's style-bias analysis is from 2024 (updated June 2025), about two years old, not three. The episode does not establish whether anyone has re-run it. - Whether the one-number indexes use model-only or adapter results for each test is an open question; the pages read do not say. - Tracker figures move. As of mid-September, Epoch's index had Astra at 166 and Claude Fable 5.1 at 164, and Artificial Analysis had the two tied at 53 (different scales). Both put Astra at or near the top. - The zero-to-97% harness example in ARC Prize's paper is one environment and one model (Claude Opus 4.6), not a general rule. - OpenAI's launch page says Astra "saturates ARC-AGI-3 with a 99.9% score" and quotes ARC Prize's Greg Kamradt saying Astra is "effectively reaching human parity on the benchmark." It does not claim AGI. - No data shows which lab gives METR the most free tokens. CHAPTERS 00:00 Same Model, Two Scores: The Cold Open 00:32 Who Graded It? The Problem With AI Test Scores 03:28 Disclosure: Ainsley Runs on Claude 03:50 Same Model, Two Scores: 62.7% vs 99.9% 04:49 What a Provider Adapter Actually Does 08:08 Four Questions for Every AI Score 10:23 The "How" Problem Behind the Number 12:14 One-Number Indexes: Do They Settle It? 14:58 Follow the Money: Who Funds the Scoreboards 18:07 The Sliced-Bread Problem and METR's Disclosures 21:25 What Gets Showcased: Arena's 2024 Style Analysis 24:02 Does Disclosure Earn Trust? 27:00 Who Checks the Checkers? METR and Free Model Access 30:19 Why It Matters and the Homework Subscribe for new episodes every Monday and Wednesday: Apple Podcasts, YouTube, Spotify. #SurvivingAI #AIBenchmarks #ARCAGI3 #AI Please visit our website for more information - Surviving AI: Navigate the Future






