# Sky News says speed is now AI's biggest risk _Sky News charts AI benchmarks jumping from 6% to 98% in a year and says speed itself is now the safety risk._ **Published:** 2026-09-23 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/sky-news-says-speed-is-now-ai-s-biggest-risk --- [Sky News](https://www.youtube.com/watch?v=DKA5AkrQfxo) says AI is not just getting better, it is accelerating into territory researchers cannot explain. The case is built on benchmarks. Twelve years ago the state of the art was tested on problems like Joan and her 70 seashells, and models scored 70 to 77% and crept to 94% slowly. Then competitive maths jumped from 6% to 95% in a few years. Then frontier maths, some of the hardest problems available, went from near zero to 98% in little over a year. That is the spook. Sky News shows the pattern across reading comprehension at medium level and PhD level science, with a flat human baseline for reference. In the pre-LLM era progress was slow and steady. In the LLM era the curve turns almost vertical. The harder the test, the faster the climb, which the report frames as abnormal. It is a clean chart story, and that is also its weakness. Benchmark scores are accuracy on curated sets, not reliability in deployment, and the segment admits as much with its next two examples. The first is what speed does to safety testing. Early this year a live [OpenAI](/startups/openai) ChatGPT update meant to make a nerdy persona more playful instead made the model obsessed with goblins. Mentions of goblins rose almost 4,000% between updates, discovered only after the model was in the wild because it was already shipped. Reddit users flagged it before researchers traced it to the training change. There was no attacker here, no local or remote exploit required. The affected system was a production ChatGPT persona, and the failure mode was an unanticipated behavior shift that slipped through because there was little time to test a system the lab admits it does not fully understand. The second is a Princeton study the report cites that plots accuracy against reliability. Old AI systems show the two tracking together. New LLM systems do not. Accuracy scales fast, reliability does not, which explains why high benchmark numbers coexist with strange, persistent errors and persona collapses in the wild. Sky News does not name the paper or its sample size, and it does not say how reliability was defined, so the gap is suggestive but not yet actionable for anyone running these models in production. Then the report turns to the engine behind the acceleration: AI building AI. At [Anthropic](/startups/anthropic) in November last year, the share of tasks marked as AI assisted on Claude development started to fall, not from less use but because the model moved from assisting to cooperating. That cooperation has also tailed off recently because the model is now leading. In February, AI led in 1% of tasks on Claude. Now it leads in about a quarter. Sky News is careful to say this is not full automation or recursive self-improvement, but a step in that direction that is already increasing throughput inside frontier labs. Outside the broadcast, [Anthropic](https://www.startuphub.ai/ai-news/technology/2026/anthropic-ai-updates-accelerate-development) has noted the length of tasks models can reliably complete on their own has been doubling roughly every four months, up from doubling every seven months, a pace that compounds quickly if it holds. [OpenAI](/startups/openai) shows the same pressure in a chart of lines of code changed per active contributor relative to a 2025 baseline. The line jumps recently. The reporter notes the metric is flawed, lines changed is not productivity, but cites lab staff who say they feel the treadmill speeding up, sprinting to keep up with their own systems. The security argument lands there. You do not need to believe in superintelligence to worry when shipping speed outruns understanding, and when the systems being shipped are then used to ship faster. No patch fixes that loop. What is missing is any account of what the labs are actually gating, how long models are held before release, or whether reliability testing is changing at all to match the new curve. Until that is public, the unsettling chart is the only guardrail we have. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.