Progress
We've seen some exciting developments in the past few weeks:
- OpenAI announced they had used their latest internal AI model, Astra, to make ten advances in mathematics and theoretical computer science. Mathematician Levent Alpöge used Anthropic's Fable 5 model to find a counterexample to the 87-year-old Jacobian Conjecture.
- Intelligence is becoming much cheaper; there have been significant advances on the price-performance frontier. For example, OpenAI cut prices for GPT-5.6 Luna by 80% and will give unlimited access to Luna for free users. I had my Opus 4.5 moment back in November. To illustrate how big this is: looking at benchmarks, I'd estimate Luna is comparable to or better than Opus 4.5, at more than 20x lower API prices. And it's not just OpenAI pushing the frontier: Chinese open-weights model DeepSeek V4 Flash is doing the same.
- ARC-AGI 3 is an interactive reasoning benchmark designed to measure human-like intelligence in AI agents. Prime Intellect announced surpassing the human-expert baseline, using an agent harness to score 95.5%.
But technological progress is not the same as actual progress.