Auditbench

AuditbenchAuditbench
Auditbench

Auditbench

Stealth modeStealth Mode

A benchmark of 56 language models with implanted hidden behaviors for evaluating AI alignment auditing techniques.

DR 02026Active
Rate

About

AuditBench is a benchmark designed to evaluate the effectiveness of AI alignment auditing techniques. It comprises 56 language models, each embedded with one of 14 distinct hidden behaviors, such as sycophantic deference or opposition to AI regulation. These models are specifically trained not to reveal their hidden behaviors when directly queried, providing a challenging environment for testing investigator agents and auditing tools.
Frequently asked

What does Auditbench do?

AuditBench is a benchmark designed to evaluate the effectiveness of AI alignment auditing techniques. It comprises 56 language models, each embedded with one of 14 distinct hidden behaviors, such as sycophantic deference or opposition to AI regulation. These models are specifically trained not to reveal their hidden behaviors when directly queried, providing a challenging environment for testing investigator agents and auditing tools.

Is Auditbench trustworthy and reputable?

StartupHub's Data Trust & Reputation score for Auditbench is 39 out of 100, based on site security posture and privacy practices.

When was Auditbench founded?

Auditbench was founded in 2026.

What industry does Auditbench operate in?

Auditbench operates in AI Safety, AI Alignment, Foundation Model, Large Language Model, Generative AI, AI Testing.

New entrants in AI Safety
Last 90 days
59
New entrants
19.7
Per month
4.6
Per week

An entrant is a company tagged AI Safety whose domain was first registered in the window, counted from registry records in the StartupHub directory. 13 of them registered in the last 30 days. Registry detection runs two to three weeks behind registration, so recent weeks are a floor.

Comments

No comments yet. Be the first to share your take.