THE CLOUD AGENT BENCHMARK v0.2

Real cloud tasks.
Proven capability.

An evaluation of AI agents on cloud operations. Reproducible tasks, isolated infrastructure, and verifiable outcomes.

POWERED BY VERA EMULATOR

The cloud environment behind CloudOpsBench. Agents work with emulated AWS resources in isolated, inspectable environments.

Explore Vera Emulator ↗
THE EVALUATION CYCLEFIG. 01
Same conditions. Observable outcomes.
THE RESULTS

Leaderboard

● Production
2 model / agent pairsv0.2AWS / emulatedSelect any row to explore
Pass · 6Fail · 6Timeout · 0Error · 12
Tasks by model and harness. Each colored square links to one recorded trial.
Taskclaude-opus-5-5claude-codegpt-6-astracodex
aws/promote-lambda-live-alias
aws/tighten-kinesis-resource-policy
aws/tighten-sqs-redrive-allow-policy

One square per trial. Click for details; hover for usage. — means no recorded trials. Multiple runs stay separate.

How outcomes are classified

This is trial history, not just the latest summary. Orange means the agent actually hit its execution timeout and did not pass; setup, verifier, infrastructure and emulator failures are red. A long duration alone is not a timeout. Gray means the grader returned a failing result without an execution error.

Resolution rate = passed / valid attempts. Evaluation errors excluded.2 entries / v0.2