THE CLOUD AGENT BENCHMARK v0.2
Real cloud tasks.
Proven capability.
An evaluation of AI agents on cloud operations. Reproducible tasks, isolated infrastructure, and verifiable outcomes.

POWERED BY VERA EMULATOR
The cloud environment behind CloudOpsBench. Agents work with emulated AWS resources in isolated, inspectable environments.
Explore Vera Emulator ↗THE EVALUATION CYCLEFIG. 01
THE RESULTS
Leaderboard
2 model / agent pairsv0.2AWS / emulatedSelect any row to explore
One square per trial. Click for details; hover for usage. — means no recorded trials. Multiple runs stay separate.
How outcomes are classified
This is trial history, not just the latest summary. Orange means the agent actually hit its execution timeout and did not pass; setup, verifier, infrastructure and emulator failures are red. A long duration alone is not a timeout. Gray means the grader returned a failing result without an execution error.
Resolution rate = passed / valid attempts. Evaluation errors excluded.2 entries / v0.2
What’s behind a score?How the benchmark works ↗