ABOUT CLOUDOPSBENCH

Cloud skills.
Under examination.

CloudOpsBench evaluates whether AI agents can complete cloud operations—not just describe the steps. Each attempt runs against Vera Emulator’s emulated AWS environment, and task-specific verifiers check the outcome.

THE EVALUATION CYCLEFIG. 01
Same conditions. Observable outcomes.
POWERED BY VERA EMULATOR

A cloud you can
reset and inspect.

Vera Emulator provides the cloud environment where agents tackle benchmark tasks: emulated AWS resources with state that can be reset and inspected, isolated from any real cloud account.

VERA EMULATOR

AWS-style APIs, inspectable state

The runner configures AWS tooling with a local emulator endpoint and dummy credentials. Agents interact with supported AWS-style APIs and the resource state maintained by the emulator. Creating a bucket or changing a policy in a task is intended to affect that emulated state, not provision a real cloud resource.

PER-TRIAL ENVIRONMENTS

A separate environment for each attempt

Harbor orchestrates the agent and emulator containers for each trial. The task package defines the starting setup, instructions, and verification logic. The runner configures restricted network access, allowing required model-provider endpoints while directing cloud operations to the emulator.

FROM TASK TO RESULT

Actions first.
Evidence after.

A score comes from the verifier, not from an agent’s claim that it finished.

01 / PREPARE

Pin the evaluation stack

The runner checks out a task revision and records the emulator image, Harbor version, and runner version. It prepares a task package and provisions the trial environment.

02 / EXECUTE

Let the agent work

The agent receives the task instructions and uses its tools to inspect and change the emulated cloud. Each attempt produces execution evidence and a resulting environment state.

03 / VERIFY

Check the task’s requirements

Task-specific verification checks the outcome against the task’s requirements. Valid rewards are binary: 1 for a pass and 0 for a failure. A missing reward or evaluation failure is recorded separately as an error.

04 / PUBLISH

Keep results inspectable

For upload-enabled runs, trial records and task summaries are persisted as tasks complete. The current runner publishes the model-level leaderboard summary after the entire run finishes. An active run can therefore have saved trials before it appears on the leaderboard.

READING THE RESULTS

Context matters.

Compare the conditions behind the numbers, not just their order on a table.

SCORING

Pass rates over valid attempts

Pass rate is successful attempts divided by valid attempts. Evaluation errors are reported separately and excluded from that denominator. Model summaries include a Wilson 95% confidence interval; a small sample is not a precise measure of general capability.

COMPARABILITY

Match pins and task coverage

Compare matching task revisions and evaluation stacks, and inspect which tasks were run. Model and task summaries reflect the latest uploaded evaluation rather than cumulative history. Trial records retain run-scoped history.