Worlds without a 3D artist
Describe a scenario; a frontier world model builds the place: splats, a collider mesh and, on the standard tier, a metric scale. Built once, reused for free.
Worlds to test in, generated by a frontier 3D world model. Compute that runs on our pools or on your own machines. Campaigns of thousands of scenarios, retried, deduplicated and indexed. And, if you want it, a measure of whether your tests would notice your robot getting worse.
Runs on
Each tile is one task: a scenario, a world, a verdict. The dispatcher hands them out, a managed pool or your own machines run them, and every verdict lands in the index with whatever decided it.
24 tasks, one world each, simulated in this pageFAIL: a time to collision under 1.2 s
You build the robot. We handle the places it is tested in, the machines that run the tests, and the record of what happened.
Describe a scenario; a frontier world model builds the place: splats, a collider mesh and, on the standard tier, a metric scale. Built once, reused for free.
Managed pools run each task on RunPod Serverless and scale to zero when idle. We are rolling these out now.
When a test needs your simulator, your robot's stack, a GPU or data that must not leave the lab, one runner binary pulls work over HTTPS. Nothing connects in.
No model judges your robot. A verdict comes from your checker or an expression you wrote, and says whether it could have failed at all.
What it does not do yet, said plainly: a generated world is static, with no dynamics or moving actors, so it tests perception, not closed-loop control. And whether a score in a generated world predicts one on the bench or the road has not been measured.
Three moves from a JSON file to thousands of verdicts, with nothing to operate in between.
# nightly.json
{ "pool": "lab-cpu", "dedupe": true,
"tasks": [{ "spec": {
"world": {
"scenario": "cut-in.xosc",
"model": "draft" } } }, ...] }A pool, a priority, and up to 100,000 task specs per request. Every submission is quoted before it is accepted, and scenarios from the catalogue become signed inputs.
$ runner --pool=lab-cpu \
--dispatcher=$SIMCLOUD_URL \
--token=$SIMCLOUD_TOKEN
leased c_7q2f.0.118 attempt 1
heartbeat ok
completed succeeded 41.2 sA managed pool, where each task runs on RunPod Serverless and nothing idles. Or your own machines: one runner binary that pulls leases, runs each task as a local process, and reports back.
Leases with tokens and TTLs, retries with backoff, parking after repeated failures, fair scheduling between teams, and a queryable result index.
On a self-hosted pool a task is any command that reads a spec and writes a run summary. Exit 75 to ask for a retry; everything else is yours. Managed pools run the world-model engine. Ask for the task contract.
Durable rows, fair leases, admission control and a queryable index, all on the edge near your runners.
One task, five states
Every task is a row before it is a queue entry. Leases carry a token and a deadline, heartbeats extend them, and a runner that dies simply hands its work back.
Pool lab-gpu
A campaign joins the pool at its current virtual time, so a 100,000-task sweep never starves the ten-task smoke test submitted after it - and priority still jumps the queue when you need it to.
maxInflight 50 · 31 leased
Cap in-flight tasks so a single campaign cannot flood a GPU pool. Releases are asynchronous, so the cap costs nothing on the hot path.
Re-submitted, dedupe on
Turn on dedupe and re-submitting last night's campaign only runs the scenarios that changed or failed. Keys are content hashes of the canonical spec.
cut-in.xosc
Upload .xosc files, get strict validation and an engine-neutral JSON model with entities, triggers and actions your tooling can reason about.
GET /v1/results?where= spec.map=rotterdam, output.summary.metrics.collisions>0 &state=succeeded
Every terminal task lands in an index with its spec, output and duration. Filter with a small expression language and page through millions of rows.
Everything a regression pipeline needs, without a platform team to run it.
On a self-hosted pool, any process that reads a spec and writes a summary. No SDK inside it is required.
Managed RunPod Serverless pools, or a single runner binary on your own machines.
Workspaces can pin campaign state to the Asia-Pacific region.
Counters for every campaign in the console, polling less as a campaign goes quiet.
Idempotency keys on submissions; duplicate tasks are rejected, not run twice.
Durations, exit codes, outputs and specs, queryable per tenant.
Content-addressed uploads, tags, search and signed download links.
Usage by day, pool or campaign, priced per provider credit; a world served from cache is free.
Credentials stay where the compute is, never in a task spec. On a self-hosted pool the control plane sees only specs and summaries.
On your own runners, only task specs and summaries cross the boundary; artefacts stay where you put them.
Runner tokens can be limited to named pools; user tokens cannot manage the workspace.
Inputs are fetched with HMAC-signed, expiring URLs, never bearer tokens.
HttpOnly, same-site cookies with cross-site writes refused at the edge.
Every credit and charge is a ledger row, and a charge comes from the provider's own receipt.
Workspaces can pin state to the Asia-Pacific region.
Optional add-on
A thousand passing scenarios prove nothing if none of them would fail. Mutational injects faults into the system under test, sweeps your corpus, and says which scenarios actually guard against which failures.
A separate product, with its own account and pricing. SimCloud works without it.
Which faults each scenario catches, through which check, and which scenarios catch nothing at all.
A worklist of the gaps, most urgent first, and which scenarios are safe to remove.
Did this build lose coverage? Asked of a shared history, with an exit code your pipeline understands.
Run it on your own machines today. In preview, a whole sweep runs as one task on a SimCloud managed pool.
Mutational and SimCloud are both made by Mutational. Each has its own account; neither needs the other.
Free credit on every new workspace, enough for a few dozen draft worlds, and unlimited runs inside them.