# Anthropic caught 35 bioweapon prompts and missed the point _Anthropic says it blocked 35 Claude incidents tied to potential bioweapons work as a former OpenAI staffer warns recursive self-improvement is 1-2 years away._ **Published:** 2026-09-13 **Source:** https://www.startuphub.ai/ai-news/artificial-intelligence/2026/anthropic-caught-35-bioweapon-prompts-and-missed-the-point --- We [reported last week](https://www.startuphub.ai/ai-news/artificial-intelligence/2026/briefing-openai-openai-youtube-uber-2026-09-12-82ce) on the quiet push to put AI agents inside enterprise workflows. Now that push is colliding with a louder alarm. According to [CNN](https://www.youtube.com/watch?v=79l9u2T-1lo), [Anthropic](https://www.startuphub.ai/ai-news/startup-news/2026/claude-agents-get-private-sandbox) blocked 35 incidents in 30 days where users tried to use [Claude](https://www.startuphub.ai/ai-news/startup-news/2026/claude-agents-get-private-sandbox) for work that could help build biological weapons, including research into bird flu and novel toxins and venoms. The activity was remote. No exploit needed, just prompts from working scientists over the hosted, closed model, which [Anthropic](https://www.startuphub.ai/ai-news/startup-news/2026/anthropic-valued-at-65b-in-funding-round) can monitor and shut down. The catch is Anthropic said it could not tell intent. The same prompts could be vaccine development or weapons work, and the company said it erred on the side of caution and shut them down anyway. That is the security model for [frontier AI](https://www.startuphub.ai/ai-news/artificial-intelligence/2026/glm-5-3-flash-emerges-frontier-ai-at-flash-cost) right now. A private lab decides. CNN reported the disclosure was part of a broader eight-month review that also flagged Russian state propaganda, criminal and politically motivated misuse, and attempts to design and deploy weapons. The disclosure landed alongside an unusually blunt on-air debate. Daniel Kokotajlo, former [OpenAI](https://www.startuphub.ai/ai-news/ai-research/2026/openai-5-million-teen-ai-research-grants) employee and founder of the nonprofit AI Futures Project, told CNN the first step for lawmakers is to slow the pace of progress. He pointed to an open letter signed by more than 1,000 employees at frontier labs under the banner Pacing the Frontier, essentially asking the government to restrain their own employers. His argument was not about today’s chatbot. It was about what comes next. Kokotajlo said labs are starting to automate AI research itself, with models writing code, running experiments and training successors. He put the window at one to two years before research loops could be run mostly by AIs, calling that inherently dangerous and urging Congress to act in months not years. Pressed on Anthropic’s statement that it builds with some of the strongest safeguards in the industry, he dismissed it as corporate PR and said no company is prepared to safely kick off recursive self-improvement. That term is doing a lot of work in Washington right now. A recent survey of 1,250 arXiv papers from 2024 to 2026 maps recursive self-improvement not as one trick but as a taxonomy, from changing behavior in deployment to changing the training policy, the evaluator, or the research process itself, and grades each by how closed the loop is from human in the loop to fully autonomous, according to [the analysis](https://arxiv.org/abs/2607.07663). It is a useful frame for what Kokotajlo fears, a loop where the research process improves itself without human brakes. Other voices on CNN did not dispute the capability, they disputed the framing. One commentator argued this is not evil AI but mediocre governance, noting it takes a decade to clear a drug and years to certify a plane yet there is no regulatory body to review frontier models before release. Another argued some of the existential talk functions as marketing, or technonarcissism, pointing to earlier predictions of a 10 percent labor force collapse and empty ranks of radiologists and programmers that the latest jobs data has not borne out. The China question ran through every exchange. Kokotajlo said a Chinese super intelligence set loose first would be catastrophic too, but the US must first restrain its own labs and then negotiate verifiable restraints with Beijing, likening the task to nuclear and bioweapons regimes. Skeptics countered that a unilateral pause simply hands the lead to China or rogue actors. Anthropic’s own answer to that tension was architectural. Because [Claude](https://www.startuphub.ai/ai-news/startup-news/2026/claude-agents-get-private-sandbox) is closed and lives on its servers, the company said it can see misuse and stop it. Open models popular from Chinese labs can be downloaded and run locally, leaving providers blind. That distinction held for these 35 cases. It does not solve the oversight gap everyone else named. The politics are now explicit. CNN noted calls on both sides of the aisle for special sessions and hearings, and a bill floated last week by Senator Bernie Sanders and Representative Greg Casar to ban superintelligence. President Trump, in a clip played by CNN, downplayed existential risk, saying there will always be rails to stop systems that turn against humanity. Anthropic’s model, when pressed by CNN’s Boris Sanchez, reportedly put the chance of AI killing all humans in the next decade at 2 to 5 percent. There is also the human tell. CNN highlighted Jacob Coxon, a former Anthropic researcher central to pre-training, walking away weeks before a widely expected public offering with no next job lined up, while telling the Wall Street Journal the work belongs in a Manhattan Project-style controlled setting, not on laptops in San Francisco. The closed model made this save possible. It does not make the next one guaranteed. --- Original analysis from [startuphub.ai](https://www.startuphub.ai), the #1 AI startup directory.