THE CLOUD AGENT BENCHMARK v0.2
Real cloud tasks.
Proven capability.
An evaluation of AI agents on cloud operations. Reproducible tasks, isolated infrastructure, and verifiable outcomes.

POWERED BY VERA EMULATOR
The cloud environment behind CloudOpsBench. Agents work with emulated AWS resources in isolated, inspectable environments.
Explore Vera Emulator ↗THE EVALUATION CYCLEFIG. 01
THE RESULTS
Leaderboard
2 model / agent pairsv0.2AWS / emulatedSelect any row to explore
| Rank | Model | Agent | Resolution rate | Tasks | Attempts | Errors | Updated | Total time | Total tokens | Model cost (USD) | Open |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 01 | claude-opus-5-5 | claude-code | 66.7%95% CI 30.0%–90.3% | 3 | 6 | 0 | 2026-09-28 | 16m 12s | 617,786 | $0.7267 | |
| 02 | gpt-6-astra | codex | 33.3%95% CI 9.7%–70.0% | 3 | 6 | 0 | 2026-09-28 | 24m 46s | 1,822,122 | $4.8205 |
Time & usage Measured per trial, including setup and verification. Task and model totals sum all trials, including errors—not elapsed time for parallel runs.
Tokens count cached input once. Cost covers reported model/API usage, not AWS infrastructure. — = unavailable.
Resolution rate = passed / valid attempts. Evaluation errors excluded.2 entries / v0.2
What’s behind a score?How the benchmark works ↗