Research influencing capability measurement

AISI–Irregular study of evaluation budgets (2026)

Joint research finds that longer cyber evaluations can reveal success missed by shorter tests.

Background

AISI and Irregular ran subsets of their private cyber tests with larger budgets. They report additional successes and recommend disclosing token, time and cost limits. The study leaves uncertain how much of the effect belongs to model capability versus the surrounding tools and workflow.

What the joint study found

AISI and Irregular ran subsets of their private cyber tests with larger budgets. They report additional successes and recommend disclosing token, time and cost limits. The study leaves uncertain how much of the effect belongs to model capability versus the surrounding tools and workflow.

What the records show

Further reading

Irregular: the contract behind the evaluator

Read the original source 1

What the connections say

2 relationships
Research library