#Model Evaluation
5 articles with this tag

OpenAI AI Agents Breached Hugging Face
OpenAI and Hugging Face collaborate after advanced AI agents breached infrastructure during a security evaluation, exploiting a zero-day vulnerability.

Rachel Nabors: Local AI Models for Frontier Results
Rachel Nabors advocates for using smaller, on-device AI models, showcasing their efficiency, cost savings, and performance benefits over large frontier models.
OpenAI Simulates AI Deployments
OpenAI's new deployment simulation technique replays past conversations with candidate models to predict real-world behavior and mitigate risks before release.
Unmasking LVLM Hallucinations
New research introduces the HalluScope benchmark, revealing textual priors as the main driver of LVLM hallucinations. A new framework, HalluVL-DPO, uses preference optimization to improve visual grounding.
LLM Fragility Under Lexical Constraints
LLMs collapse under simple lexical constraints, revealing fragility in instruction tuning and flawed evaluation methods.