CloudOpsBench evaluates whether AI agents can complete cloud operations—not just describe the steps. Each attempt runs against Vera Emulator’s emulated AWS environment, and task-specific verifiers check the outcome.
Vera Emulator provides the cloud environment where agents tackle benchmark tasks: emulated AWS resources with state that can be reset and inspected, isolated from any real cloud account.
VERA EMULATOR
AWS-style APIs, inspectable state
The runner configures AWS tooling with a local emulator endpoint and dummy credentials. Agents interact with supported AWS-style APIs and the resource state maintained by the emulator. Creating a bucket or changing a policy in a task is intended to affect that emulated state, not provision a real cloud resource.
PER-TRIAL ENVIRONMENTS
A separate environment for each attempt
Harbor orchestrates the agent and emulator containers for each trial. The task package defines the starting setup, instructions, and verification logic. The runner configures restricted network access, allowing required model-provider endpoints while directing cloud operations to the emulator.
FROM TASK TO RESULT
Actions first. Evidence after.
A score comes from the verifier, not from an agent’s claim that it finished.
01 / PREPARE
Pin the evaluation stack
The runner checks out a task revision and records the emulator image, Harbor version, and runner version. It prepares a task package and provisions the trial environment.
02 / EXECUTE
Let the agent work
The agent receives the task instructions and uses its tools to inspect and change the emulated cloud. Each attempt produces execution evidence and a resulting environment state.
03 / VERIFY
Check the task’s requirements
Task-specific verification checks the outcome against the task’s requirements. Valid rewards are binary: 1 for a pass and 0 for a failure. A missing reward or evaluation failure is recorded separately as an error.
04 / PUBLISH
Keep results inspectable
For upload-enabled runs, trial records and task summaries are persisted as tasks complete. The current runner publishes the model-level leaderboard summary after the entire run finishes. An active run can therefore have saved trials before it appears on the leaderboard.
READING THE RESULTS
Context matters.
Compare the conditions behind the numbers, not just their order on a table.
SCORING
Pass rates over valid attempts
Pass rate is successful attempts divided by valid attempts. Evaluation errors are reported separately and excluded from that denominator. Model summaries include a Wilson 95% confidence interval; a small sample is not a precise measure of general capability.
COMPARABILITY
Match pins and task coverage
Compare matching task revisions and evaluation stacks, and inspect which tasks were run. Model and task summaries reflect the latest uploaded evaluation rather than cumulative history. Trial records retain run-scoped history.