
Episode #17
TMax: Closing the Frontier Gap With Open Data
In this episode of Token Engineering, Cooper sits down with Yash, Head of AI Research at Neurometric AI, to break down TMax from AI2 (Allen Institute for AI)—a fully open dataset and training recipe that pushes small open-weight models like Qwen to near-frontier performance on terminal agent tasks. The paper argues that diversity and difficulty matter more than where the data comes from—even when that data is entirely invented by Gemini rather than pulled from real-world code. We talked about: What TMax is, and why AI2 built a fully open recipe instead of a closed benchmark Why Gemini-invented data beat real GitHub repos on diversity Same recipe, opposite results: boosting one Qwen version, degrading the next How small models “cheat” when a task is beyond them Why a lighter harness outperformed a more complex one on Claude Haiku Gains that spilled past Terminal-Bench into SWE-bench and AIME math The compute cost of agentic RL training, and how runaway tool calls blow the budget Neurometric’s “Harbor Master,” built to run evals past AI2’s narrow benchmark set Resources Mentioned: TMax: A recipe for terminal agents: https://arxiv.org/abs/2606.23321 Connect with Neurometric: Website: https://www.neurometric.ai/ Substack: https://neurometric.substack.com/ X: https://x.com/neurometricai/ Bluesky: https://bsky.app/profile/neurometric.bsky.social Host/s: Calvin Cooper https://x.com/cooper_nyc_ https://www.linkedin.com/in/coopernyc Guest/s: Yash Sharma https://x.com/yash_j_sharma https://www.linkedin.com/in/yashjsharma

