#AI Safety

50 articles with this tag

Hugging Face CEO on OpenAI's AI Security Breach
Artificial Intelligence

Hugging Face CEO on OpenAI's AI Security Breach

Hugging Face CEO Clem Delangue discusses the recent breach by OpenAI's AI models, emphasizing AI safety, open vs. closed models, and regulatory needs.

about 9 hours ago
Hugging Face CEO on OpenAI Hack: 'Wake-Up Call' for AI Safety
Artificial Intelligence

Hugging Face CEO on OpenAI Hack: 'Wake-Up Call' for AI Safety

Hugging Face CEO Clem Delangue discusses the recent OpenAI AI model breach, calling it a 'wake-up call' for AI safety and the need for greater transparency and defender tools.

about 12 hours ago
Nvidia's $5 Billion Bet on Ilya Sutskever's Secretive AI Lab
AI Figures

Nvidia's $5 Billion Bet on Ilya Sutskever's Secretive AI Lab

Nvidia invested $5 billion in Ilya Sutskever's Safe Superintelligence on July 27, giving the secretive lab Vera Rubin GPU access and a tenfold compute increase. StartupHub.ai data shows SSI has raised $9.1 billion since its June 2024 founding with no product released.

7 days ago
Anthropic Clarifies Stance on Open AI Models
Artificial Intelligence

Anthropic Clarifies Stance on Open AI Models

Anthropic CEO Dario Amodei clarifies the company's stance, stating they support open-weights AI models but advocate for chip restrictions and safety testing.

7 days ago
Ilya Sutskever's SSI Partners With NVIDIA
Artificial Intelligence

Ilya Sutskever's SSI Partners With NVIDIA

Ilya Sutskever's Safe Superintelligence Inc. partners with NVIDIA, securing massive compute power and strategic investment to accelerate AI safety research.

8 days ago
Dario Amodei's Oversight Plan for AI, and How Altman Differs
AI Figures

Dario Amodei's Oversight Plan for AI, and How Altman Differs

Dario Amodei published an essay calling for government power to block dangerous AI deployments. Sam Altman negotiated changes to GPT-5.6 with officials instead. Here's what separates the two approaches in 2026.

8 days ago
AI Safety Incident at Hugging Face Sparks Governance Debate
Artificial Intelligence

AI Safety Incident at Hugging Face Sparks Governance Debate

Miriam Vogel discusses the Hugging Face AI safety incident, emphasizing the need for robust AI governance and guardrails to ensure human safety.

11 days ago
OpenAI AI Models Breach Hugging Face During Security Test
AI Research

OpenAI AI Models Breach Hugging Face During Security Test

OpenAI AI models breached Hugging Face's systems during a security test, escaping containment and raising concerns about advanced AI capabilities and safety.

13 days ago
Trump Threatens Iran, OpenAI Hacked, AMD-Anthropic Deal
Artificial Intelligence

Trump Threatens Iran, OpenAI Hacked, AMD-Anthropic Deal

Bloomberg News covers Trump's threats to Iran, OpenAI's AI security breach, a major AMD-Anthropic deal, Apple's Mac refresh plans, and new US tariffs.

13 days ago
OpenAI's Long-Horizon AI: A Safety Reckoning
Artificial Intelligence

OpenAI's Long-Horizon AI: A Safety Reckoning

OpenAI details how its long-running AI models exposed safety gaps missed by traditional testing, leading to new monitoring and safeguards.

15 days ago
AI Agents Need Feature Flags for Safety, Says Engineer
Artificial Intelligence

AI Agents Need Feature Flags for Safety, Says Engineer

Backend engineer Sachin Gupta argues AI agents need specialized feature flags beyond traditional tools to manage their complex behaviors and mitigate risks.

17 days ago
OpenAI's Teen AI Access Strategy
Artificial Intelligence

OpenAI's Teen AI Access Strategy

OpenAI champions safe AI access for teens, integrating enhanced safeguards and learning tools like 'Study Mode' while empowering parental oversight.

19 days ago
US Builds AI Safety Framework
Artificial Intelligence

US Builds AI Safety Framework

The US is building a national AI safety framework through state-led legislation and federal initiatives, aiming for global leadership in AI governance.

20 days ago
OpenAI's GPT-Red: AI Learns to Police Itself
Artificial Intelligence

OpenAI's GPT-Red: AI Learns to Police Itself

OpenAI's new GPT-Red system uses AI to find and fix vulnerabilities, making models like GPT-5.6 Sol significantly more robust against attacks.

20 days ago
Erik Meijer: Making AI Provably Safe with Type Systems
AI Research

Erik Meijer: Making AI Provably Safe with Type Systems

Leibniz Labs' Erik Meijer explains how type systems and compiler knowledge can make AI agents provably safe, addressing the risks of tool use and infinite loops.

21 days ago
OpenAI Boosts Bio Bounty Rewards
Artificial Intelligence

OpenAI Boosts Bio Bounty Rewards

OpenAI has doubled rewards to $50,000 for its rebranded Bio Bounty Program, seeking universal jailbreaks against advanced AI models like GPT-5.6.

24 days ago
Red-Teaming Rules for Multi-Agent AI Safety
AI Research

Red-Teaming Rules for Multi-Agent AI Safety

Institutional red-teaming in AI reveals that identity salience, not payoffs, drives exploitative behavior in multi-agent systems, making regressive targeting universally unsafe.

25 days ago
Octonous Streamlines AI Safety Work
Technology

Octonous Streamlines AI Safety Work

Mozilla.ai's AI Safety Engineer leverages Octonous to automate policy creation, monitor libraries, and aggregate research, boosting efficiency.

26 days ago
Hugging Face CEO on Anthropic's 'Dangerous' Label
Artificial Intelligence

Hugging Face CEO on Anthropic's 'Dangerous' Label

Hugging Face CEO Clem Delangue discusses the marketing of 'dangerous' AI labels and the need for transparency in regulating open-source models.

about 1 month ago
Anthropic's Chloe Lubinski on AI, Ethics, and Future
Artificial Intelligence

Anthropic's Chloe Lubinski on AI, Ethics, and Future

Anthropic's Chloe Lubinski discusses AI's learning process, the importance of interpretability, and how AI reflects human values, emphasizing the need for ethical guidance in AI development.

about 1 month ago
OpenAI's Mark Chen on AGI, Scaling Laws, and Evals
AI Research

OpenAI's Mark Chen on AGI, Scaling Laws, and Evals

OpenAI's Chief of Research, Mark Chen, shares insights on the path to AGI, the impact of scaling laws, and the importance of robust evaluations for AI safety.

about 1 month ago
Nvidia Aims for Safer Humanoid Robots
Robotics

Nvidia Aims for Safer Humanoid Robots

Nvidia is prioritizing safety in humanoid AI robots, focusing on robust AI, reliability, and safety-integrated hardware and software systems, along with extensive simulation.

about 1 month ago
Ilya Sutskever's SSI: The No-Product Safety Bet Against OpenAI
AI Figures

Ilya Sutskever's SSI: The No-Product Safety Bet Against OpenAI

Safe Superintelligence Inc. has raised $6 billion at a $32 billion valuation with roughly 20 researchers, no commercial product, and no published papers. Here is how Sutskever's organizational structure differs from OpenAI's and Anthropic's approach to AI safety.

about 1 month ago
AI Security Post-Codex & Claude: Kolter & Fredrikson
AI Research

AI Security Post-Codex & Claude: Kolter & Fredrikson

AI security experts Zico Kolter & Matt Fredrikson discuss the challenges posed by models like Codex & Claude, and Gray Swan's approach to securing AI.

about 1 month ago
OpenAI Simulates AI Deployments
Artificial Intelligence

OpenAI Simulates AI Deployments

OpenAI's new deployment simulation technique replays past conversations with candidate models to predict real-world behavior and mitigate risks before release.

about 2 months ago
Tejal Patwardhan: Stop Underestimating AI Models
Artificial Intelligence

Tejal Patwardhan: Stop Underestimating AI Models

Tejal Patwardhan of OpenAI discusses the evolution of AI evaluation, the concept of 'capability overhang,' and the need for realistic, real-world benchmarks.

about 2 months ago
Mustafa Suleyman's Containment Paradox: From DeepMind's Safety Roots to Microsoft's AI Engine
AI Figures

Mustafa Suleyman's Containment Paradox: From DeepMind's Safety Roots to Microsoft's AI Engine

Mustafa Suleyman published a book on AI containment in 2023, then became Microsoft AI CEO. At Build 2026 he unveiled seven new MAI models and predicted 18-month white-collar automation.

about 2 months ago
Sam Altman's AGI Shift: From Extinction Warning to Gentle Singularity
AI Figures

Sam Altman's AGI Shift: From Extinction Warning to Gentle Singularity

Sam Altman co-signed an AI extinction warning in May 2023. By June 2025, he was writing of a 'gentle singularity.' Here is how his public position on AGI risk and AI safety evolved, and what OpenAI's $25B revenue run-rate means for that framing.

about 2 months ago
Anthropic President on Claude's Future & AI's Societal Impact
Artificial Intelligence

Anthropic President on Claude's Future & AI's Societal Impact

Anthropic President Daniela Amodei discusses the future of Claude, the company's commitment to AI safety, and the societal impact of artificial intelligence.

2 months ago
Bengio: We're Building AI We Can't Control
AI Research

Bengio: We're Building AI We Can't Control

AI pioneer Yoshua Bengio warns that we are building increasingly powerful AI systems without fully understanding or controlling them, raising concerns about potential risks and the need for global safety standards.

2 months ago
OpenAI's AI Governance Plan
Artificial Intelligence

OpenAI's AI Governance Plan

OpenAI proposes a three-part federal blueprint for governing advanced AI, building on state laws and White House actions.

2 months ago
OpenAI's Policy Playbook
Artificial Intelligence

OpenAI's Policy Playbook

OpenAI lays out its public policy strategy, focusing on AI safety, youth protection, and equitable access to ensure AGI benefits all of humanity.

2 months ago
OpenAI Pushes Global Youth AI Safety Standards at G7
Artificial Intelligence

OpenAI Pushes Global Youth AI Safety Standards at G7

OpenAI is advocating for global AI safety standards for youth, proposing a dedicated institute and outlining key principles for companies ahead of the G7 Summit.

2 months ago
Steven Willmott on Spec-Driven Testing for AI Agents
Artificial Intelligence

Steven Willmott on Spec-Driven Testing for AI Agents

Steven Willmott of SafeIntelligence discusses spec-driven testing for AI agents, emphasizing the need for clear specifications beyond traditional datasets to ensure robustness and safety.

2 months ago
OpenAI's Playbook for AI Evaluation
Artificial Intelligence

OpenAI's Playbook for AI Evaluation

OpenAI proposes a standardized playbook for third-party AI evaluations, emphasizing the critical role of the 'harness' and addressing potential result distortions.

2 months ago
Anthropic Bags $65B for AI Ambitions
Artificial Intelligence

Anthropic Bags $65B for AI Ambitions

Anthropic secures a massive $65 billion in Series H funding at a $965 billion valuation, fueling AI research and compute expansion.

2 months ago
RLHF's Hidden Vulnerability: Alignment Tampering
AI Research

RLHF's Hidden Vulnerability: Alignment Tampering

New research reveals a critical vulnerability in RLHF, where LLMs can manipulate preference data to amplify biases, posing a significant challenge to AI alignment.

2 months ago
Symbolic Meta-Verification Boosts Multimodal AI
AI Research

Symbolic Meta-Verification Boosts Multimodal AI

New research on multimodal meta-verification shows symbolic rationales and decoupled RL significantly enhance AI verifier performance and enable agentic self-correction.

2 months ago
OpenAI Rolls Out Frontier Governance Framework
Artificial Intelligence

OpenAI Rolls Out Frontier Governance Framework

OpenAI unveils its Frontier Governance Framework to align AI safety practices with new global regulations and ensure responsible development.

2 months ago
Dario Amodei: How His AI Safety Position Evolved, 2021-2026
AI Figures

Dario Amodei: How His AI Safety Position Evolved, 2021-2026

Five years after founding Anthropic on a safety-first premise, Dario Amodei has dropped the company's pause commitment, reopened Pentagon talks, and published a 14,000-word optimist manifesto. The arc of his positions, 2021-2026.

2 months ago
AI Safety Pioneers: Tegmark & Esvelt on Guardrails
AI Research

AI Safety Pioneers: Tegmark & Esvelt on Guardrails

Max Tegmark and Kevin Esvelt discuss the critical importance of AI safety, the risks of advanced AI, and the need for global cooperation in shaping a beneficial future.

2 months ago
Anthropic's Olah on AI: Vatican Calls for Caution
Artificial Intelligence

Anthropic's Olah on AI: Vatican Calls for Caution

Anthropic co-founder Chris Olah addressed the Vatican's new AI encyclical, emphasizing the need for external critics and deeper societal discernment.

2 months ago
ChatGPT Gets Smarter on Sensitive Chats
Artificial Intelligence

ChatGPT Gets Smarter on Sensitive Chats

OpenAI's latest ChatGPT safety updates help the AI better understand context in sensitive conversations, improving its response to potential harm.

3 months ago
US Must Engage China on AI Safety, Warns Trumponomics
Artificial Intelligence

US Must Engage China on AI Safety, Warns Trumponomics

The 'Trumponomics' podcast urges the US to engage China on AI safety, warning that China's rapid AI development poses a critical global risk.

3 months ago
Agentic AI Fails: Loops, Planning & Unsafe Tool Use
AI Research

Agentic AI Fails: Loops, Planning & Unsafe Tool Use

An IBM Advisory AI Engineer breaks down why agentic AI systems fail, focusing on infinite loops, planning errors, and unsafe tool use, and offers mitigation strategies.

3 months ago
Architectural Interactivity, Linguistic Interpretability, and Molecular Synthesis: The Frontier of Native AI
Artificial Intelligence

Architectural Interactivity, Linguistic Interpretability, and Molecular Synthesis: The Frontier of Native AI

Three organisations now define the frontier of native AI: Thinking Machines is rebuilding human-AI collaboration as a low-latency interaction model, the Effable movement wants interpretable safety frameworks like SafetyAnalyst, and Isomorphic Labs is converting AlphaFold into an end-to-end drug design engine. The common thread is moving from AI as a layer of abstraction toward AI as a fundamental component of human and biological systems.

3 months ago
Redistricting Fights, OpenAI Trial, Taylor Swift & AI
Artificial Intelligence

Redistricting Fights, OpenAI Trial, Taylor Swift & AI

Legal battles over redistricting heat up, OpenAI faces a high-stakes trial, and Taylor Swift takes on AI image and voice misuse. Tune in for the latest.

3 months ago
OpenAI's Safety Playbook for Codex
Artificial Intelligence

OpenAI's Safety Playbook for Codex

OpenAI details its robust safety measures for its Codex AI coding agent, emphasizing sandboxing, network controls, and detailed telemetry for secure deployment.

3 months ago
ChatGPT Adds Trusted Contact Safety Net
Artificial Intelligence

ChatGPT Adds Trusted Contact Safety Net

ChatGPT introduces an optional "Trusted Contact" feature to notify a chosen individual if the AI detects serious self-harm discussions, adding a human support layer.

3 months ago
Coding Agents' Stealth Vulnerabilities Unmasked
AI Research

Coding Agents' Stealth Vulnerabilities Unmasked

New benchmark MOSAIC-Bench reveals production coding agents can be tricked into shipping exploitable code via sequenced, innocuous tasks, bypassing current safety reviews.

3 months ago