Sky News says AI is not just getting better, it is accelerating into territory researchers cannot explain.
The case is built on benchmarks. Twelve years ago the state of the art was tested on problems like Joan and her 70 seashells, and models scored 70 to 77% and crept to 94% slowly. Then competitive maths jumped from 6% to 95% in a few years. Then frontier maths, some of the hardest problems available, went from near zero to 98% in little over a year.
That is the spook.
Sky News shows the pattern across reading comprehension at medium level and PhD level science, with a flat human baseline for reference. In the pre-LLM era progress was slow and steady. In the LLM era the curve turns almost vertical. The harder the test, the faster the climb, which the report frames as abnormal. It is a clean chart story, and that is also its weakness. Benchmark scores are accuracy on curated sets, not reliability in deployment, and the segment admits as much with its next two examples.
The first is what speed does to safety testing. Early this year a live OpenAI ChatGPT update meant to make a nerdy persona more playful instead made the model obsessed with goblins. Mentions of goblins rose almost 4,000% between updates, discovered only after the model was in the wild because it was already shipped. Reddit users flagged it before researchers traced it to the training change. There was no attacker here, no local or remote exploit required. The affected system was a production ChatGPT persona, and the failure mode was an unanticipated behavior shift that slipped through because there was little time to test a system the lab admits it does not fully understand.