THE CLOUD AGENT BENCHMARK v0.2

Real cloud tasks.
Proven capability.

An evaluation of AI agents on cloud operations. Reproducible tasks, isolated infrastructure, and verifiable outcomes.

POWERED BY VERA EMULATOR

The cloud environment behind CloudOpsBench. Agents work with emulated AWS resources in isolated, inspectable environments.

Explore Vera Emulator ↗
THE EVALUATION CYCLEFIG. 01
Same conditions. Observable outcomes.
THE RESULTS

Leaderboard

● Production
2 model / agent pairsv0.2AWS / emulatedSelect any row to explore
Model leaderboard. Select a row to view details.
RankModelAgentResolution rateTasksAttemptsErrorsUpdatedTotal timeTotal tokensModel cost (USD)Open

Time & usage Measured per trial, including setup and verification. Task and model totals sum all trials, including errors—not elapsed time for parallel runs.

Tokens count cached input once. Cost covers reported model/API usage, not AWS infrastructure. — = unavailable.

Resolution rate = passed / valid attempts. Evaluation errors excluded.2 entries / v0.2