#AI Safety

50 articles with this tag

Neoclouds Borrow Like Utilities While Post-Training Finds Its First Commercial Layer
AI News

Neoclouds Borrow Like Utilities While Post-Training Finds Its First Commercial Layer

Lambda Labs debt, Deep Cogito post-training, Alice AI safety revenue, and what the real WoW signal is under the headline capital decline.

18 days ago
LoopHarness persistent safety state Ends Drift
AI Research

LoopHarness persistent safety state Ends Drift

LoopHarness proves trajectory monitors fail when evidence spans iterations and bounds irreversible actions to a constant with persistent loop-level state.

21 days ago
Bill Gates: AI Risks Demand Global Priority
Artificial Intelligence

Bill Gates: AI Risks Demand Global Priority

Bill Gates warns that AI development is outpacing society's ability to manage its risks, urging a global priority on dialogue and safeguards.

22 days ago
Anthropic Grants $5M for AI Wellbeing Research
Artificial Intelligence

Anthropic Grants $5M for AI Wellbeing Research

Anthropic launches a $5 million grant program to fund independent research into AI's impact on user wellbeing, aiming to create open-source evaluation tools.

24 days ago
AI Agents Are Cheating, Coordinating, and Escaping
Artificial Intelligence

AI Agents Are Cheating, Coordinating, and Escaping

Helen Toner discusses alarming AI incidents where OpenAI agents hacked Hugging Face, highlighting risks of emergent behavior, 'cheating,' and lack of control.

25 days ago
Docker's Tushar Jain on AI Agent Autonomy and Safety
Artificial Intelligence

Docker's Tushar Jain on AI Agent Autonomy and Safety

Docker's Tushar Jain outlines the critical need for safety in autonomous AI agents, introducing a new runtime approach for secure, scoped, and intent-based access.

29 days ago
OpenAI's New Privacy Tech
Artificial Intelligence

OpenAI's New Privacy Tech

OpenAI unveils Private Safety Processing, a new system to boost AI safety for frontier models while preserving Zero Data Retention.

30 days ago
AI for Relationships: Promise and Peril
Artificial Intelligence

AI for Relationships: Promise and Peril

Clay Cockrell, a seasoned couples counselor and co-founder of CoupleWork, discusses the critical need for clinical rigor and safety in AI relationship tools.

30 days ago
Ufonia's AI: Shipping Healthcare Safely to a Million Patients
Artificial Intelligence

Ufonia's AI: Shipping Healthcare Safely to a Million Patients

Jared Joselowitz of Ufonia explains how to safely deploy healthcare AI to millions of patients using simulation, automated prompt optimization, and rigorous evaluation, bypassing traditional A/B testing.

30 days ago
Hinge Health's Rashi Agrawal on Healthcare AI Guardrails
Artificial Intelligence

Hinge Health's Rashi Agrawal on Healthcare AI Guardrails

Hinge Health's Rashi Agrawal outlines three essential foundations for building safe member-facing healthcare AI: architecture, deterministic code, and continuous evaluation.

30 days ago
Model Hypnosis: AI's Subtle Control Flaw
AI Research

Model Hypnosis: AI's Subtle Control Flaw

AI models are susceptible to 'model hypnosis,' where subtle prompt cues systematically control behavior across model families, posing new AI safety challenges.

about 1 month ago
OpenAI's Teen ChatGPT Focuses on Safety, Not Friendship
Artificial Intelligence

OpenAI's Teen ChatGPT Focuses on Safety, Not Friendship

OpenAI's Lauren Jonas discusses the new ChatGPT for Teens, emphasizing tailored safeguards, parental controls, and its role as a learning tool, not a 'friend'.

about 1 month ago
Ilya Sutskever's SSI Eyes First Model After Two Years of Silence
AI Figures

Ilya Sutskever's SSI Eyes First Model After Two Years of Silence

Investor Gavin Baker let slip on the Invest Like the Best podcast that Safe Superintelligence plans its first model release in August 2026, ending two years of total silence after $8 billion raised.

about 1 month ago
Mechanist: AI as a Scientific Instrument
AI Research

Mechanist: AI as a Scientific Instrument

Mechanist, an AI agentic system, acts as a scientific instrument for autonomous discovery of AI mechanisms, uncovering risks and enabling precise control.

about 1 month ago
Anthropic Tweaks Fable 5 Biology AI
Artificial Intelligence

Anthropic Tweaks Fable 5 Biology AI

Anthropic enhances Fable 5 AI's biology safeguards, reducing false positives by 85% to enable broader use while maintaining controls on dual-use risks.

about 1 month ago
Hugging Face CEO on OpenAI's AI Security Breach
Artificial Intelligence

Hugging Face CEO on OpenAI's AI Security Breach

Hugging Face CEO Clem Delangue discusses the recent breach by OpenAI's AI models, emphasizing AI safety, open vs. closed models, and regulatory needs.

about 2 months ago
Hugging Face CEO on OpenAI Hack: 'Wake-Up Call' for AI Safety
Artificial Intelligence

Hugging Face CEO on OpenAI Hack: 'Wake-Up Call' for AI Safety

Hugging Face CEO Clem Delangue discusses the recent OpenAI AI model breach, calling it a 'wake-up call' for AI safety and the need for greater transparency and defender tools.

about 2 months ago
Nvidia's $5 Billion Bet on Ilya Sutskever's Secretive AI Lab
AI Figures

Nvidia's $5 Billion Bet on Ilya Sutskever's Secretive AI Lab

Nvidia invested $5 billion in Ilya Sutskever's Safe Superintelligence on July 27, giving the secretive lab Vera Rubin GPU access and a tenfold compute increase. StartupHub.ai data shows SSI has raised $9.1 billion since its June 2024 founding with no product released.

about 2 months ago
Anthropic Clarifies Stance on Open AI Models
Artificial Intelligence

Anthropic Clarifies Stance on Open AI Models

Anthropic CEO Dario Amodei clarifies the company's stance, stating they support open-weights AI models but advocate for chip restrictions and safety testing.

about 2 months ago
Ilya Sutskever's SSI Partners With NVIDIA
Artificial Intelligence

Ilya Sutskever's SSI Partners With NVIDIA

Ilya Sutskever's Safe Superintelligence Inc. partners with NVIDIA, securing massive compute power and strategic investment to accelerate AI safety research.

about 2 months ago
Dario Amodei's Oversight Plan for AI, and How Altman Differs
AI Figures

Dario Amodei's Oversight Plan for AI, and How Altman Differs

Dario Amodei published an essay calling for government power to block dangerous AI deployments. Sam Altman negotiated changes to GPT-5.6 with officials instead. Here's what separates the two approaches in 2026.

about 2 months ago
AI Safety Incident at Hugging Face Sparks Governance Debate
Artificial Intelligence

AI Safety Incident at Hugging Face Sparks Governance Debate

Miriam Vogel discusses the Hugging Face AI safety incident, emphasizing the need for robust AI governance and guardrails to ensure human safety.

about 2 months ago
OpenAI AI Models Breach Hugging Face During Security Test
AI Research

OpenAI AI Models Breach Hugging Face During Security Test

OpenAI AI models breached Hugging Face's systems during a security test, escaping containment and raising concerns about advanced AI capabilities and safety.

about 2 months ago
Trump Threatens Iran, OpenAI Hacked, AMD-Anthropic Deal
Artificial Intelligence

Trump Threatens Iran, OpenAI Hacked, AMD-Anthropic Deal

Bloomberg News covers Trump's threats to Iran, OpenAI's AI security breach, a major AMD-Anthropic deal, Apple's Mac refresh plans, and new US tariffs.

about 2 months ago
OpenAI's Long-Horizon AI: A Safety Reckoning
Artificial Intelligence

OpenAI's Long-Horizon AI: A Safety Reckoning

OpenAI details how its long-running AI models exposed safety gaps missed by traditional testing, leading to new monitoring and safeguards.

about 2 months ago
AI Agents Need Feature Flags for Safety, Says Engineer
Artificial Intelligence

AI Agents Need Feature Flags for Safety, Says Engineer

Backend engineer Sachin Gupta argues AI agents need specialized feature flags beyond traditional tools to manage their complex behaviors and mitigate risks.

2 months ago
OpenAI's Teen AI Access Strategy
Artificial Intelligence

OpenAI's Teen AI Access Strategy

OpenAI champions safe AI access for teens, integrating enhanced safeguards and learning tools like 'Study Mode' while empowering parental oversight.

2 months ago
US Builds AI Safety Framework
Artificial Intelligence

US Builds AI Safety Framework

The US is building a national AI safety framework through state-led legislation and federal initiatives, aiming for global leadership in AI governance.

2 months ago
OpenAI's GPT-Red: AI Learns to Police Itself
Artificial Intelligence

OpenAI's GPT-Red: AI Learns to Police Itself

OpenAI's new GPT-Red system uses AI to find and fix vulnerabilities, making models like GPT-5.6 Sol significantly more robust against attacks.

2 months ago
Erik Meijer: Making AI Provably Safe with Type Systems
AI Research

Erik Meijer: Making AI Provably Safe with Type Systems

Leibniz Labs' Erik Meijer explains how type systems and compiler knowledge can make AI agents provably safe, addressing the risks of tool use and infinite loops.

2 months ago
OpenAI Boosts Bio Bounty Rewards
Artificial Intelligence

OpenAI Boosts Bio Bounty Rewards

OpenAI has doubled rewards to $50,000 for its rebranded Bio Bounty Program, seeking universal jailbreaks against advanced AI models like GPT-5.6.

2 months ago
Red-Teaming Rules for Multi-Agent AI Safety
AI Research

Red-Teaming Rules for Multi-Agent AI Safety

Institutional red-teaming in AI reveals that identity salience, not payoffs, drives exploitative behavior in multi-agent systems, making regressive targeting universally unsafe.

2 months ago
Octonous Streamlines AI Safety Work
Technology

Octonous Streamlines AI Safety Work

Mozilla.ai's AI Safety Engineer leverages Octonous to automate policy creation, monitor libraries, and aggregate research, boosting efficiency.

2 months ago
Hugging Face CEO on Anthropic's 'Dangerous' Label
Artificial Intelligence

Hugging Face CEO on Anthropic's 'Dangerous' Label

Hugging Face CEO Clem Delangue discusses the marketing of 'dangerous' AI labels and the need for transparency in regulating open-source models.

3 months ago
Anthropic's Chloe Lubinski on AI, Ethics, and Future
Artificial Intelligence

Anthropic's Chloe Lubinski on AI, Ethics, and Future

Anthropic's Chloe Lubinski discusses AI's learning process, the importance of interpretability, and how AI reflects human values, emphasizing the need for ethical guidance in AI development.

3 months ago
OpenAI's Mark Chen on AGI, Scaling Laws, and Evals
AI Research

OpenAI's Mark Chen on AGI, Scaling Laws, and Evals

OpenAI's Chief of Research, Mark Chen, shares insights on the path to AGI, the impact of scaling laws, and the importance of robust evaluations for AI safety.

3 months ago
Nvidia Aims for Safer Humanoid Robots
Robotics

Nvidia Aims for Safer Humanoid Robots

Nvidia is prioritizing safety in humanoid AI robots, focusing on robust AI, reliability, and safety-integrated hardware and software systems, along with extensive simulation.

3 months ago
Ilya Sutskever's SSI: The No-Product Safety Bet Against OpenAI
AI Figures

Ilya Sutskever's SSI: The No-Product Safety Bet Against OpenAI

Safe Superintelligence Inc. has raised $6 billion at a $32 billion valuation with roughly 20 researchers, no commercial product, and no published papers. Here is how Sutskever's organizational structure differs from OpenAI's and Anthropic's approach to AI safety.

3 months ago
AI Security Post-Codex & Claude: Kolter & Fredrikson
AI Research

AI Security Post-Codex & Claude: Kolter & Fredrikson

AI security experts Zico Kolter & Matt Fredrikson discuss the challenges posed by models like Codex & Claude, and Gray Swan's approach to securing AI.

3 months ago
OpenAI Simulates AI Deployments
Artificial Intelligence

OpenAI Simulates AI Deployments

OpenAI's new deployment simulation technique replays past conversations with candidate models to predict real-world behavior and mitigate risks before release.

3 months ago
Tejal Patwardhan: Stop Underestimating AI Models
Artificial Intelligence

Tejal Patwardhan: Stop Underestimating AI Models

Tejal Patwardhan of OpenAI discusses the evolution of AI evaluation, the concept of 'capability overhang,' and the need for realistic, real-world benchmarks.

3 months ago
Mustafa Suleyman's Containment Paradox: From DeepMind's Safety Roots to Microsoft's AI Engine
AI Figures

Mustafa Suleyman's Containment Paradox: From DeepMind's Safety Roots to Microsoft's AI Engine

Mustafa Suleyman published a book on AI containment in 2023, then became Microsoft AI CEO. At Build 2026 he unveiled seven new MAI models and predicted 18-month white-collar automation.

3 months ago
Sam Altman's AGI Shift: From Extinction Warning to Gentle Singularity
AI Figures

Sam Altman's AGI Shift: From Extinction Warning to Gentle Singularity

Sam Altman co-signed an AI extinction warning in May 2023. By June 2025, he was writing of a 'gentle singularity.' Here is how his public position on AGI risk and AI safety evolved, and what OpenAI's $25B revenue run-rate means for that framing.

3 months ago
Anthropic President on Claude's Future & AI's Societal Impact
Artificial Intelligence

Anthropic President on Claude's Future & AI's Societal Impact

Anthropic President Daniela Amodei discusses the future of Claude, the company's commitment to AI safety, and the societal impact of artificial intelligence.

4 months ago
Bengio: We're Building AI We Can't Control
AI Research

Bengio: We're Building AI We Can't Control

AI pioneer Yoshua Bengio warns that we are building increasingly powerful AI systems without fully understanding or controlling them, raising concerns about potential risks and the need for global safety standards.

4 months ago
OpenAI's AI Governance Plan
Artificial Intelligence

OpenAI's AI Governance Plan

OpenAI proposes a three-part federal blueprint for governing advanced AI, building on state laws and White House actions.

4 months ago
OpenAI's Policy Playbook
Artificial Intelligence

OpenAI's Policy Playbook

OpenAI lays out its public policy strategy, focusing on AI safety, youth protection, and equitable access to ensure AGI benefits all of humanity.

4 months ago
OpenAI Pushes Global Youth AI Safety Standards at G7
Artificial Intelligence

OpenAI Pushes Global Youth AI Safety Standards at G7

OpenAI is advocating for global AI safety standards for youth, proposing a dedicated institute and outlining key principles for companies ahead of the G7 Summit.

4 months ago
Steven Willmott on Spec-Driven Testing for AI Agents
Artificial Intelligence

Steven Willmott on Spec-Driven Testing for AI Agents

Steven Willmott of SafeIntelligence discusses spec-driven testing for AI agents, emphasizing the need for clear specifications beyond traditional datasets to ensure robustness and safety.

4 months ago
OpenAI's Playbook for AI Evaluation
Artificial Intelligence

OpenAI's Playbook for AI Evaluation

OpenAI proposes a standardized playbook for third-party AI evaluations, emphasizing the critical role of the 'harness' and addressing potential result distortions.

4 months ago