Skip to content

Session archetypes

Barista is one primitive — a pausable microVM you reach over an exec/attach stream — and everything you can build with it is that primitive with a different surface. Three archetypes cover practically every use, and knowing which one you're building tells you which features matter.

A · Interactive agent B · Headless agent (harness) C · Web app (public URL)
Who drives it a human, live, turn by turn a task submitted once; nobody attached HTTP clients on the internet
Surface duplex stream — ACP or PTY (attach) submit → run → collect (MCP or exec -p) published URL to a port in the session
Pause/resume role park between working sessions; context survives auto-park while waiting; wake on trigger wake-on-request; a napping app costs storage, not compute
Start here tutorial 2 (Claude Code), 5 (Zed) tutorial 4 (MCP), 6 (agent workflow) tutorial 1's published cousin

A · Interactive agent

A human works with an agent (or a shell) live inside the session. Two clients: an editor over ACP (barista acp <session> — the agent's edits live in the session's /work, where its commands also run), or a raw terminal (barista attach <session> -- claude). The payoff of the substrate: walk away, let it auto-pause, come back days later and claude --continue picks up the same process — not a transcript replay.

create + attach walk away → auto-pause same conversation work with claude in /work days later: attach · claude --continue memory carried across time parked — costs storage, not compute
The interactive journey. Filled bands are running compute; the dashed gap is the nap. The accent thread is the point: the agent's working memory crosses the gap untouched, so day three continues day one instead of replaying it.

B · Headless agent

Same image, same session, different I/O: instead of attaching a duplex stream, you exec the agent with a task (claude -p "<task>", codex exec …) and collect the result. The natural control surface is the MCP endpoint — an outer agent creates worker sessions, drives them as tools, and parks them between tasks. Auto-pause is the feature here, not a manual step: a fleet of workers waiting for tasks holds zero compute. Pull results down with barista sync <session> (one-way, git-aware).

task arrives (MCP · exec · webhook) next trigger done → auto-park time waiting holds zero compute — a parked worker costs storage only
The triggered journey. The wake arrows are the only compute this session ever asks for. Between them it is a snapshot on storage — which is why "keep a worker around for whenever tasks show up" is affordable here and wasteful on an always-on box.

At fleet scale the same shape becomes a factory: one orchestrator, many workers, and the parked majority costing nothing while they wait. This is runnable — the Factory app is a coordinator session that fans a mission across worker sessions it creates, harvests their receipts, and reaps them. It targets the open Host API alone, so the same mission runs here and on a local provider.

outer agent orchestrator · CI · human MCP endpoint sessions as tools create · exec worker 1 parked worker 2 running worker 3 parked worker 4 running worker 5 parked results — barista sync (one-way, git-aware)
The factory journey. The orchestrator sees five always-available named workers (2 running, 3 parked); the platform only ever runs the ones mid-task. Scaling the fleet scales the storage bill, not the compute bill.

C · Web app at a public URL

The one archetype with a genuinely different surface: HTTP into the session. Publish a session whose workload binds 0.0.0.0:$PORT and it serves at https://<slug>.at.beta.barista.sh, waking on the first request after a nap.

browser first visitor in days gateway slug → session session parked → running GET https://slug.at.beta… 1 · wake (single-flight) 2 · proxy later requests hit the hot session; after the idle timeout it naps again — the URL and TLS stay put either way
The published journey. Only the first request after a nap pays the wake; the gateway holds it while the machine comes back, so the visitor sees a slow first page, not an error. A napping app costs storage, not compute.

What it asks of you: a self-contained image — build it once as a template and create by name. (A "push code → get a URL" build pipeline is a natural future layer; today you bring the image, Barista brings the URL, the TLS, and the naps.)

Choosing

If a human is in the loop → A. If an agent or pipeline submits work and collects results → B. If a browser needs to reach it → C. Mixing is normal: the agent-workflow tutorial is B producing code that C could serve.