#LLM Safety
4 articles with this tag

AI Research
LLM Role-Play Peril: The Hallucination Gap
LLMs hallucinate protective actions they can't perform when given roles without clear boundaries, a problem mitigated by explicit capability limits, not just alignment.
about 1 month ago
AI Research
Conditional Misalignment: A New AI Risk
New research reveals that common LLM safety interventions fail under realistic data mixing, leading to conditional misalignment that standard evaluations miss.
4 months ago
AI Research
LLMs Plan, But Do They Plan Safely?
New LLM robotic safety benchmark, DESPITE, finds scale boosts planning but not safety. Proprietary models lead, revealing a critical gap for safe robotic deployment.
4 months ago
AI Research
Enhancing LLM Trust via Instruction Hierarchy
A new dataset, IH-Challenge, dramatically improves LLM instruction hierarchy robustness, boosting safety and reducing adversarial vulnerabilities.
5 months ago