#AI Safety
50 articles with this tag

Hugging Face CEO on OpenAI's AI Security Breach
Hugging Face CEO Clem Delangue discusses the recent breach by OpenAI's AI models, emphasizing AI safety, open vs. closed models, and regulatory needs.

Hugging Face CEO on OpenAI Hack: 'Wake-Up Call' for AI Safety
Hugging Face CEO Clem Delangue discusses the recent OpenAI AI model breach, calling it a 'wake-up call' for AI safety and the need for greater transparency and defender tools.

Nvidia's $5 Billion Bet on Ilya Sutskever's Secretive AI Lab
Nvidia invested $5 billion in Ilya Sutskever's Safe Superintelligence on July 27, giving the secretive lab Vera Rubin GPU access and a tenfold compute increase. StartupHub.ai data shows SSI has raised $9.1 billion since its June 2024 founding with no product released.

Anthropic Clarifies Stance on Open AI Models
Anthropic CEO Dario Amodei clarifies the company's stance, stating they support open-weights AI models but advocate for chip restrictions and safety testing.

Ilya Sutskever's SSI Partners With NVIDIA
Ilya Sutskever's Safe Superintelligence Inc. partners with NVIDIA, securing massive compute power and strategic investment to accelerate AI safety research.

Dario Amodei's Oversight Plan for AI, and How Altman Differs
Dario Amodei published an essay calling for government power to block dangerous AI deployments. Sam Altman negotiated changes to GPT-5.6 with officials instead. Here's what separates the two approaches in 2026.

AI Safety Incident at Hugging Face Sparks Governance Debate
Miriam Vogel discusses the Hugging Face AI safety incident, emphasizing the need for robust AI governance and guardrails to ensure human safety.

OpenAI AI Models Breach Hugging Face During Security Test
OpenAI AI models breached Hugging Face's systems during a security test, escaping containment and raising concerns about advanced AI capabilities and safety.

Trump Threatens Iran, OpenAI Hacked, AMD-Anthropic Deal
Bloomberg News covers Trump's threats to Iran, OpenAI's AI security breach, a major AMD-Anthropic deal, Apple's Mac refresh plans, and new US tariffs.

OpenAI's Long-Horizon AI: A Safety Reckoning
OpenAI details how its long-running AI models exposed safety gaps missed by traditional testing, leading to new monitoring and safeguards.

AI Agents Need Feature Flags for Safety, Says Engineer
Backend engineer Sachin Gupta argues AI agents need specialized feature flags beyond traditional tools to manage their complex behaviors and mitigate risks.

OpenAI's Teen AI Access Strategy
OpenAI champions safe AI access for teens, integrating enhanced safeguards and learning tools like 'Study Mode' while empowering parental oversight.

US Builds AI Safety Framework
The US is building a national AI safety framework through state-led legislation and federal initiatives, aiming for global leadership in AI governance.

OpenAI's GPT-Red: AI Learns to Police Itself
OpenAI's new GPT-Red system uses AI to find and fix vulnerabilities, making models like GPT-5.6 Sol significantly more robust against attacks.

Erik Meijer: Making AI Provably Safe with Type Systems
Leibniz Labs' Erik Meijer explains how type systems and compiler knowledge can make AI agents provably safe, addressing the risks of tool use and infinite loops.

OpenAI Boosts Bio Bounty Rewards
OpenAI has doubled rewards to $50,000 for its rebranded Bio Bounty Program, seeking universal jailbreaks against advanced AI models like GPT-5.6.

Red-Teaming Rules for Multi-Agent AI Safety
Institutional red-teaming in AI reveals that identity salience, not payoffs, drives exploitative behavior in multi-agent systems, making regressive targeting universally unsafe.

Octonous Streamlines AI Safety Work
Mozilla.ai's AI Safety Engineer leverages Octonous to automate policy creation, monitor libraries, and aggregate research, boosting efficiency.

Hugging Face CEO on Anthropic's 'Dangerous' Label
Hugging Face CEO Clem Delangue discusses the marketing of 'dangerous' AI labels and the need for transparency in regulating open-source models.

Anthropic's Chloe Lubinski on AI, Ethics, and Future
Anthropic's Chloe Lubinski discusses AI's learning process, the importance of interpretability, and how AI reflects human values, emphasizing the need for ethical guidance in AI development.

OpenAI's Mark Chen on AGI, Scaling Laws, and Evals
OpenAI's Chief of Research, Mark Chen, shares insights on the path to AGI, the impact of scaling laws, and the importance of robust evaluations for AI safety.

Nvidia Aims for Safer Humanoid Robots
Nvidia is prioritizing safety in humanoid AI robots, focusing on robust AI, reliability, and safety-integrated hardware and software systems, along with extensive simulation.

Ilya Sutskever's SSI: The No-Product Safety Bet Against OpenAI
Safe Superintelligence Inc. has raised $6 billion at a $32 billion valuation with roughly 20 researchers, no commercial product, and no published papers. Here is how Sutskever's organizational structure differs from OpenAI's and Anthropic's approach to AI safety.

AI Security Post-Codex & Claude: Kolter & Fredrikson
AI security experts Zico Kolter & Matt Fredrikson discuss the challenges posed by models like Codex & Claude, and Gray Swan's approach to securing AI.
OpenAI Simulates AI Deployments
OpenAI's new deployment simulation technique replays past conversations with candidate models to predict real-world behavior and mitigate risks before release.

Tejal Patwardhan: Stop Underestimating AI Models
Tejal Patwardhan of OpenAI discusses the evolution of AI evaluation, the concept of 'capability overhang,' and the need for realistic, real-world benchmarks.

Mustafa Suleyman's Containment Paradox: From DeepMind's Safety Roots to Microsoft's AI Engine
Mustafa Suleyman published a book on AI containment in 2023, then became Microsoft AI CEO. At Build 2026 he unveiled seven new MAI models and predicted 18-month white-collar automation.

Sam Altman's AGI Shift: From Extinction Warning to Gentle Singularity
Sam Altman co-signed an AI extinction warning in May 2023. By June 2025, he was writing of a 'gentle singularity.' Here is how his public position on AGI risk and AI safety evolved, and what OpenAI's $25B revenue run-rate means for that framing.

Anthropic President on Claude's Future & AI's Societal Impact
Anthropic President Daniela Amodei discusses the future of Claude, the company's commitment to AI safety, and the societal impact of artificial intelligence.

Bengio: We're Building AI We Can't Control
AI pioneer Yoshua Bengio warns that we are building increasingly powerful AI systems without fully understanding or controlling them, raising concerns about potential risks and the need for global safety standards.
OpenAI's AI Governance Plan
OpenAI proposes a three-part federal blueprint for governing advanced AI, building on state laws and White House actions.
OpenAI's Policy Playbook
OpenAI lays out its public policy strategy, focusing on AI safety, youth protection, and equitable access to ensure AGI benefits all of humanity.
OpenAI Pushes Global Youth AI Safety Standards at G7
OpenAI is advocating for global AI safety standards for youth, proposing a dedicated institute and outlining key principles for companies ahead of the G7 Summit.

Steven Willmott on Spec-Driven Testing for AI Agents
Steven Willmott of SafeIntelligence discusses spec-driven testing for AI agents, emphasizing the need for clear specifications beyond traditional datasets to ensure robustness and safety.
OpenAI's Playbook for AI Evaluation
OpenAI proposes a standardized playbook for third-party AI evaluations, emphasizing the critical role of the 'harness' and addressing potential result distortions.

Anthropic Bags $65B for AI Ambitions
Anthropic secures a massive $65 billion in Series H funding at a $965 billion valuation, fueling AI research and compute expansion.
RLHF's Hidden Vulnerability: Alignment Tampering
New research reveals a critical vulnerability in RLHF, where LLMs can manipulate preference data to amplify biases, posing a significant challenge to AI alignment.
Symbolic Meta-Verification Boosts Multimodal AI
New research on multimodal meta-verification shows symbolic rationales and decoupled RL significantly enhance AI verifier performance and enable agentic self-correction.
OpenAI Rolls Out Frontier Governance Framework
OpenAI unveils its Frontier Governance Framework to align AI safety practices with new global regulations and ensure responsible development.

Dario Amodei: How His AI Safety Position Evolved, 2021-2026
Five years after founding Anthropic on a safety-first premise, Dario Amodei has dropped the company's pause commitment, reopened Pentagon talks, and published a 14,000-word optimist manifesto. The arc of his positions, 2021-2026.

AI Safety Pioneers: Tegmark & Esvelt on Guardrails
Max Tegmark and Kevin Esvelt discuss the critical importance of AI safety, the risks of advanced AI, and the need for global cooperation in shaping a beneficial future.

Anthropic's Olah on AI: Vatican Calls for Caution
Anthropic co-founder Chris Olah addressed the Vatican's new AI encyclical, emphasizing the need for external critics and deeper societal discernment.

ChatGPT Gets Smarter on Sensitive Chats
OpenAI's latest ChatGPT safety updates help the AI better understand context in sensitive conversations, improving its response to potential harm.

US Must Engage China on AI Safety, Warns Trumponomics
The 'Trumponomics' podcast urges the US to engage China on AI safety, warning that China's rapid AI development poses a critical global risk.

Agentic AI Fails: Loops, Planning & Unsafe Tool Use
An IBM Advisory AI Engineer breaks down why agentic AI systems fail, focusing on infinite loops, planning errors, and unsafe tool use, and offers mitigation strategies.
Architectural Interactivity, Linguistic Interpretability, and Molecular Synthesis: The Frontier of Native AI
Three organisations now define the frontier of native AI: Thinking Machines is rebuilding human-AI collaboration as a low-latency interaction model, the Effable movement wants interpretable safety frameworks like SafetyAnalyst, and Isomorphic Labs is converting AlphaFold into an end-to-end drug design engine. The common thread is moving from AI as a layer of abstraction toward AI as a fundamental component of human and biological systems.

Redistricting Fights, OpenAI Trial, Taylor Swift & AI
Legal battles over redistricting heat up, OpenAI faces a high-stakes trial, and Taylor Swift takes on AI image and voice misuse. Tune in for the latest.

OpenAI's Safety Playbook for Codex
OpenAI details its robust safety measures for its Codex AI coding agent, emphasizing sandboxing, network controls, and detailed telemetry for secure deployment.

ChatGPT Adds Trusted Contact Safety Net
ChatGPT introduces an optional "Trusted Contact" feature to notify a chosen individual if the AI detects serious self-harm discussions, adding a human support layer.

Coding Agents' Stealth Vulnerabilities Unmasked
New benchmark MOSAIC-Bench reveals production coding agents can be tricked into shipping exploitable code via sequenced, innocuous tasks, bypassing current safety reviews.