THE CLOUD AGENT BENCHMARK v0.2

Real cloud tasks.
Proven capability.

An evaluation of AI agents on cloud operations. Reproducible tasks, isolated infrastructure, and verifiable outcomes.

POWERED BY VERA EMULATOR

The cloud environment behind CloudOpsBench. Agents work with emulated AWS resources in isolated, inspectable environments.

Explore Vera Emulator ↗
THE EVALUATION CYCLEFIG. 01
Same conditions. Observable outcomes.
THE RESULTS

Leaderboard

● Production
2 model / agent pairsv0.2AWS / emulatedSelect any row to explore
Higher resolution, lower usage is better ↖
Each numbered point is a model and agent pair. Blue points are non-dominated within matching task sets and evaluation pins. Exact values and model links follow the chart.0%20%40%60%80%100%$0.0000$0.2209$0.4419$0.6628$0.8838Model cost (USD / trial)Resolution rate12
  1. claude-opus-5-5 / claude-code66.7% · $0.1211 / trial · frontier
  2. gpt-6-astra / codex33.3% · $0.8034 / trial

Usage is averaged over all recorded trials in each summary, including errors. Time includes setup and verification. Error bars show the resolution rate’s 95% confidence interval.

Frontiers compare only models with matching evaluation pins and exact task sets; different groups are not ranked against each other. At least two comparable models are required. Trial counts and error rates may differ—inspect the table before drawing conclusions.

Resolution rate = passed / valid attempts. Evaluation errors excluded.2 entries / v0.2