this spring, two things happened in the same week, and the discourse filed them under one heading.
a frontier model disproved a conjecture Paul Erdős posed in 1946. eighty years open. closed without an outline, without hints, by drawing a connection between algebraic number theory and plane geometry that no mathematician had drawn. the proof was checked and co-signed by nine mathematicians, including the same skeptics who caught OpenAI overstating an earlier result.
the same week, a benchmark called nanoGPT-Bench reported what happens when you point the best coding agents at a real machine learning research problem and let them run alone. no hints, no internet, a fixed compute budget. the best agent recovered nine percent of the progress human researchers had made over five months. a sister benchmark had already handed agents pseudocode and paper-like descriptions of every known improvement. even then they recovered less than half.
both got called AI doing science. they are not the same kind of event.
one is compression: a known space, walked faster than any human could walk it. the other is discovery: ground that had no map. mistaking one for the other changes what you think is inevitable, what you think is defensible, and what you think is still worth a human's judgment.
i. execution versus discovery: the four rungs that separate grinding a known space from reaching one nobody had mapped.
ii. the verifier is the world: why some fields are falling to AI and others are stuck, and the two incompatible bets being placed on the difference.