Skip to content

How it works

Barista Cloud is a thin, honest control plane over the open-source Barista engine: the engine owns the hard part (microVMs that pause with their memory and resume as the same live process), the cloud layer owns tenancy, access, and the surfaces you actually touch — CLI, SDK, MCP, ACP, console, and public URLs.

The pieces

CLI · SDK · console · MCP · ACP public URL slug.at.beta.barista.sh gateway auth · tenancy · quotas · templates publishing · wake-on-request coordination bucket desired/ specs · leases publish slugs barista node — hypeman + KVM runs OCI images as microVMs · forwards workload ports pauses and resumes with memory intact HTTPS · API key / OAuth wake-on-request create → desired/ reconciles: desired → microVM stamps lease (instance · addr) Contract A — mTLS gRPC: exec · attach · pause · resume · proxy
Two paths, one session. Creation is bucket-mediated (gateway writes intent, the node reconciles); live operations go straight to the node over Contract A. A request to a published URL takes the accent path: the gateway wakes the session if it's parked, then proxies.
  • Create is bucket-mediated. The gateway validates and admits your create (template resolution, size, idle policy, quota), then writes a desired/<tenant>__<name> record. The node's reconciler picks it up, materialises the microVM, and writes back a lease. Deleting the desired record releases the name.
  • Live operations go over Contract A. exec, attach, pause, resume, and status queries reach the node's gRPC API — mutually-authenticated TLS between the control plane and each node.
  • The data path is node-published. A session that serves HTTP gets a port forwarded by the node's ingress and reports a dialable address; the gateway proxies published URLs to it, validating every address against the trusted node allowlist before dialing (SSRF defense).

The substrate: a microVM is a process

Everything Barista promises follows from one unglamorous fact: to the host, a microVM is an ordinary Linux process. The virtual machine monitor (VMM) is a userspace program; the guest's RAM is a plain memory mapping inside it; the guest's disk and network are virtio devices the VMM backs with an OCI image and the host network.

one microVM = one ordinary host process (the VMM) guest RAM a plain memory mapping virtio devices disk — from the OCI image net — egress · port forward guest Linux kernel your image's processes claude · node · anything guest agent exec · attach · files host Linux kernel KVM — /dev/kvm CPU — Intel VT-x / AMD-V guest instructions run natively on the silicon
Why pause/resume is even possible. The guest's entire universe — its RAM, its devices — lives inside one host process. Capture that process's memory and device state and you have captured the machine; that snapshot (the accented block) is exactly what a pause writes out.

KVM is the piece that makes this fast. It is not an emulator: the CPU itself runs guest instructions natively in a hardware "guest mode," and the kernel's KVM module only mediates entry and exit. The VMM emulates nothing but the handful of virtio devices.

VMM (userspace) owns guest RAM · emulates virtio only KVM kernel module /dev/kvm CPU — guest mode executes the guest natively KVM_RUN enter VM exit emulate I/O the loop: run natively → exit only for device I/O → emulate → re-enter
No CPU emulation, ever. Guest code runs at native speed; the only overhead is crossing this loop when the guest touches a virtual device. That is why a microVM boots in milliseconds and runs like hardware.

Sleep and wake

Every session carries an idle timeout (your plan's default, per-session override, 0 disables). When it expires, the node captures the microVM's memory and releases its compute — the session now costs storage, not CPU or RAM. Three things wake it, transparently:

  • an exec or attach against it (wake-on-intent),
  • a request to its published URL (wake-on-request),
  • an explicit resume.
running RAM the live process state vCPUs · devices active parked snapshot RAM + disk, on storage 0 vCPU · 0 RAM resident running again the same process continues mid-thought variables · caches intact idle timeout: capture · release exec · request · resume: restore the process's memory — never rebuilt, only moved
Pause is capture-then-release. The node writes the microVM's RAM and device state to a snapshot, then frees every runtime resource. Resume maps the memory back and the guest continues — same variables, same caches, same agent conversation. A parked session is not "stopped"; it is a machine holding its breath.

Resuming restores the same process — variables, caches, an agent's conversation state — not a rebuilt container. The engine-side mechanics (snapshots, restore duties, honesty about what a runtime can guarantee) are documented on the engine site.

Two kinds of node

dev node real node
Runtime local subprocess hypeman + KVM microVM
Memory across pause/resume no yes (memory snapshot)
Allowed in production refuses to run yes

The hosted beta runs real KVM nodes; memory preservation across pause/resume is the measured, load-bearing property everything above builds on.