Skip to content

Caveman documentation

Everything you need to use Caveman. Operators should read the setup guide in the repository.

Start a build

Sign in, describe what you want on the New build screen, and press Build it. Caveman creates a project, a durable run, and queues it for a worker. You can close the browser at any time; the run continues.

Optional settings let you name a preferred stack, add constraints, record a deployment target, and set the run's budget ceiling. Everything else is inferred.

Follow a run

The run dashboard shows real state from the orchestration core: the current stage, each task and its specialist, attempts, dependencies, trusted check results, reviews, failures and recoveries, approvals, model usage and cost.

Task states: Waiting, Ready, Running, Validating, Reviewing, Revision Needed, Needs Approval, Blocked, Failed, Accepted.

What “done” means

A specialist returning work only creates a candidate. It is accepted only after every predeclared check passes in the sandbox and a fresh independent reviewer approves it — against the exact bytes submitted. The run completes only when every success criterion cites accepted work.

Checks run inside Bubblewrap with no network, a cleared environment, and resource limits. The sandbox runs Python (compile, pytest) and Node/TypeScript (the Node test runner and tsc). npm dependencies are installed by a separate isolated step with install scripts disabled, then mounted read-only. Other stacks are delivered as reviewed source and documents.

Approvals

Caveman asks only for consequential decisions: material plan changes, capability escalations, and actions such as publishing code. Each request states what, why, the risk and the exact scope. Approving binds to the scope digest you were shown; if the request changes, you are asked again.

Budgets and cost

Every run has a spending ceiling and a model-call ceiling. Caveman records provider, model, tokens, cache usage and provider-reported cost for each call. When a limit is reached the run pauses safely; raise the budget and continue if you choose.

Some providers do not report cost for every call. Caveman shows that honestly rather than estimating.

Delivery

When a run completes, Caveman assembles an archive from exactly the accepted, fingerprint-verified files plus a build report. Nothing is pushed or deployed on your behalf.

Current limitations

Live previews of generated apps are not available yet; Caveman will not render untrusted code on its own origin. Publishing to GitHub is not wired yet. Sandboxed execution supports Python and Node/TypeScript; builds and dev servers are not run. OpenRouter is the only configured model provider.

Ready? Start a build.