How it works¶
Barista Cloud is a thin, honest control plane over the open-source Barista engine: the engine owns the hard part (microVMs that pause with their memory and resume as the same live process), the cloud layer owns tenancy, access, and the surfaces you actually touch — CLI, SDK, MCP, ACP, console, and public URLs.
The pieces¶
- Create is bucket-mediated. The gateway validates and admits your create
(template resolution, size, idle policy, quota), then writes a
desired/<tenant>__<name>record. The node's reconciler picks it up, materialises the microVM, and writes back a lease. Deleting the desired record releases the name. - Live operations go over Contract A.
exec,attach,pause,resume, and status queries reach the node's gRPC API — mutually-authenticated TLS between the control plane and each node. - The data path is node-published. A session that serves HTTP gets a port forwarded by the node's ingress and reports a dialable address; the gateway proxies published URLs to it, validating every address against the trusted node allowlist before dialing (SSRF defense).
The substrate: a microVM is a process¶
Everything Barista promises follows from one unglamorous fact: to the host, a microVM is an ordinary Linux process. The virtual machine monitor (VMM) is a userspace program; the guest's RAM is a plain memory mapping inside it; the guest's disk and network are virtio devices the VMM backs with an OCI image and the host network.
KVM is the piece that makes this fast. It is not an emulator: the CPU itself runs guest instructions natively in a hardware "guest mode," and the kernel's KVM module only mediates entry and exit. The VMM emulates nothing but the handful of virtio devices.
Sleep and wake¶
Every session carries an idle timeout (your plan's default, per-session
override, 0 disables). When it expires, the node captures the microVM's
memory and releases its compute — the session now costs storage, not CPU or
RAM. Three things wake it, transparently:
- an
execorattachagainst it (wake-on-intent), - a request to its published URL (wake-on-request),
- an explicit
resume.
Resuming restores the same process — variables, caches, an agent's conversation state — not a rebuilt container. The engine-side mechanics (snapshots, restore duties, honesty about what a runtime can guarantee) are documented on the engine site.
Two kinds of node¶
| dev node | real node | |
|---|---|---|
| Runtime | local subprocess | hypeman + KVM microVM |
| Memory across pause/resume | no | yes (memory snapshot) |
| Allowed in production | refuses to run | yes |
The hosted beta runs real KVM nodes; memory preservation across pause/resume is the measured, load-bearing property everything above builds on.