Deployment Recipes¶
This guide captures reference deployment flows for running Orcheo locally during development and hosting the service for teams. Each recipe lists the required environment variables, supporting services, and common verification steps.
Local Development (PostgreSQL)¶
This setup mirrors the default configuration that the tests exercise. It is ideal when you want to iterate on nodes, run the FastAPI server, and execute LangGraph workflows from the command line.
- Install dependencies
- Configure environment variables
- Multi-workspace is always on. Users must already belong to a workspace or create one through the self-service API after login.
- Keep
ORCHEO_WORKSPACE_BACKEND=postgresandORCHEO_POSTGRES_DSNpointed at a durable database so memberships and workspace metadata survive backend restarts. - Start the API server
- Run an example workflow
- Send a websocket message to
ws://localhost:2025/ws/workflow/<workflow_id>(see the Authentication Guide for token options), or trigger a run withorcheo workflow run <workflow_id>.
Verification: Run uv run pytest to validate the environment. The test suite uses the same backend factories as the server.
Vault note: Set ORCHEO_VAULT_BACKEND=postgres and ORCHEO_VAULT_ENCRYPTION_KEY before starting the backend so credential encryption is configured from the first run.
Repository note: Local development uses the PostgreSQL workflow repository. Set ORCHEO_REPOSITORY_BACKEND=postgres and ORCHEO_POSTGRES_DSN so runs, triggers, and workflow state persist durably.
Workspace Bring-up¶
Workspace scoping is always on. Before exposing the API:
- Postgres workspace store
- Set
ORCHEO_WORKSPACE_BACKEND=postgresand provideORCHEO_POSTGRES_DSN. - Memberships
- Confirm every user has at least one workspace membership. Service tokens
and dev logins must carry
workspace_idsin their claims. - Verification
- Hit
/api/workspaces/meand confirm the Studio workspace badge to ensure the resolved workspace matches expectations.
Docker Compose (PostgreSQL, multi-container)¶
Use this recipe when you want an isolated environment that mimics production with a dedicated PostgreSQL database.
- Create
docker-compose.ymlservices: orcheo: build: . command: uvicorn orcheo_backend.app:app --host 0.0.0.0 --port 2025 environment: ORCHEO_HOST: 0.0.0.0 ORCHEO_PORT: "2025" ORCHEO_CHECKPOINT_BACKEND: postgres ORCHEO_GRAPH_STORE_BACKEND: postgres ORCHEO_REPOSITORY_BACKEND: postgres ORCHEO_WORKSPACE_BACKEND: postgres ORCHEO_CHATKIT_BACKEND: postgres ORCHEO_VAULT_BACKEND: postgres ORCHEO_VAULT_ENCRYPTION_KEY: change-me ORCHEO_POSTGRES_DSN: postgresql://orcheo:orcheo@postgres:5432/orcheo ports: - "2025:2025" depends_on: - postgres postgres: image: postgres:16 environment: POSTGRES_USER: orcheo POSTGRES_PASSWORD: orcheo POSTGRES_DB: orcheo ports: - "5432:5432" volumes: - postgres-data:/var/lib/postgresql/data volumes: postgres-data: - Build and start
- Connect
Access the API via
http://localhost:2025. The Postgres database is stored inside the named volume so runs persist across container restarts.
Verification: curl http://localhost:2025/api/system/info confirms the container is healthy.
Vault note: Rotate ORCHEO_VAULT_ENCRYPTION_KEY regularly and back up the Postgres volume alongside the database.
Lean Single Image (Backend + Studio)¶
ghcr.io/ai-colleagues/orcheo-lean packages the backend and Studio, both built
from source at the tagged revision, into one image. The backend serves Studio
on the same origin (port 2025), so there is no separate Studio container.
With Celery worker, a cron scheduler, and Redis¶
deploy/lean/docker-compose.yml runs the lean image three times (backend,
worker, and a cron scheduler) next to Redis, with in-process execution and cron
turned off on the backend. PostgreSQL is not bundled: the stack uses a Supabase database through
ORCHEO_POSTGRES_DSN, which must be set. Use the Supabase transaction pooler
connection string (Supavisor, *.pooler.supabase.com port 6543, from the
project's Connect dialog). It works over IPv4 and lets the backend, worker, and
scheduler share a few server connections. The session pooler (port 5432) holds one
server connection per client connection, so it fails with EMAXCONNSESSION
once the project's pool size (15 on small projects) is used up; if you stay on
it, lower ORCHEO_POSTGRES_POOL_MAX_SIZE. Orcheo disables server-side prepared
statements, which transaction pooling does not support.
All PostgreSQL pools check connections on checkout and enable TCP keepalives.
New connections have a 10-second connection timeout (minimum 2 seconds, matching
libpq); Linux connections also limit unacknowledged TCP data to 30 seconds. The
default pool acquisition wait is 5 seconds. max_idle is 240 seconds but only
retires connections above the pool minimum; it is not the dead-connection
safeguard. These settings are configurable through the PostgreSQL variables in
Environment Variables.
Identity SQL uses transaction-local statement (10 seconds) and lock (3 seconds)
timeouts so Supabase transaction pooling preserves the budgets. Identity calls
and workspace resolution run outside the backend event loop. They share
Starlette/AnyIO's default worker limiter (40 concurrent calls) with other synchronous
routes and dependencies. During an outage those slots can fill and delay unrelated
worker calls; pool and SQL timeouts bound individual database waits, but do not
bound time queued for a worker. Monitor worker saturation and request latency
under concurrent outage load before increasing the limiter: more threads can
increase database contention. Checkout validation also adds one database round
trip per acquisition. Orcheo requires psycopg-pool>=3.2.1 for checkout checks and
the acquisition timeout fix; psycopg[binary] supplies a supported libpq, while
source builds need libpq 12+ for tcp_user_timeout.
Database connection and timeout errors return a sanitized 503 response; email challenge issuance retains its constant response to protect account privacy. Studio preserves its session on transient refresh failures and offers an explicit retry when sign-in is unavailable. API calls retain the refresh error status and body for diagnostics; only a definitive 401 refresh rejection clears the stored session. The backend uses 401 for missing, expired, revoked, or already-rotated refresh tokens. Other statuses can come from a proxy, rate limiter, or service failure and do not prove the token is invalid. Refresh rotation is not automatically repeated after an ambiguous failure.

These limits cover different failure modes, not a single end-to-end request
deadline. A statement timeout cannot repair a network blackhole, and TCP user
timeout only bounds unacknowledged data. Validate network failure behavior on the
Linux production stack before rollout. The Linux PostgreSQL CI job also tests
packet loss using a one-second TCP user timeout and a rule restricted to the
test connection. It checks that the request fails promptly and the pool recovers.
Opt-in PostgreSQL regression tests can run against a disposable database with
ORCHEO_TEST_POSTGRES_DSN:
ORCHEO_TEST_POSTGRES_DSN=postgresql://localhost/orcheo_test uv run pytest \
tests/integration/test_postgres_resilience.py tests/identity/test_identity_concurrency.py
Set ORCHEO_TEST_TCP_PACKET_DROP=1 only on an isolated Linux runner with a local
IPv4 PostgreSQL service, iptables, and passwordless sudo to include the
packet-drop check. The test removes its connection-specific rule in finally.
The direct connection and the dedicated pooler (db.<ref>.supabase.co) are
IPv6-only. The compose file enables IPv6 on its network
(ORCHEO_LEAN_ENABLE_IPV6, default true), which needs Docker Engine 27 or
newer and a host with outbound IPv6. On Docker Desktop, pick dual IPv4/IPv6
networking under Settings > Resources > Network. Set
ORCHEO_LEAN_ENABLE_IPV6=false on older engines, where creating an IPv6
network without a configured subnet fails.
From the repository root, create deploy/lean/.env from the template and fill
it in, then start the stack:
cp deploy/lean/.env.example deploy/lean/.env
docker compose -f deploy/lean/docker-compose.yml up -d --build
Compose reads deploy/lean/.env, not the repository-root .env. --build
builds Dockerfile.lean from your checkout; the image includes the ChatKit
widgets from deploy/stack/chatkit_widgets, so nothing is mounted. To run a published
image, set ORCHEO_LEAN_IMAGE=ghcr.io/ai-colleagues/orcheo-lean:<version> and
use --no-build. ORCHEO_POSTGRES_DSN and ORCHEO_VAULT_ENCRYPTION_KEY must
be set (the optional .env is read for them). Set ORCHEO_LEAN_PUBLIC_URL to the browser-facing origin used
for invitation links, the MCP sign-in consent page and CORS (default
http://localhost:2025), and
ORCHEO_LEAN_PORT to change the host port.
Without a checkout, orcheo install --lean downloads docker-compose.yml and
.env.example from deploy/lean/ at the newest lean-v* release into
~/.orcheo/lean, prompts for the Supabase connection string and the email
domains allowed to sign in, writes .env with generated secrets and the pinned
ORCHEO_LEAN_IMAGE, and starts the stack with the published image. The
template keeps local defaults (auth disabled, CLI uploads allowed, port bound to
127.0.0.1). An https:// backend URL at the prompt switches sign-in to
required; otherwise set ORCHEO_AUTH_MODE=required before exposing it.
ORCHEO_AUTH_ALLOWED_EMAIL_DOMAINS limits sign-in to the listed domains once
sign-in is required.
Lean workflow workers default to two execution processes. Set
ORCHEO_LEAN_WORKER_CONCURRENCY to tune this against memory usage and queue
latency. The service named celery-beat runs a singleton
orcheo_backend.worker.cron_scheduler process, which checks schedules directly
and publishes their runs to Celery. Schedule checks therefore do not wait for
workflow execution slots. The lean worker ignores queued Celery cron-dispatch
tasks left over from earlier deployments (ORCHEO_CRON_DISPATCH_OWNER=scheduler).
Keep exactly one scheduler and leave
ORCHEO_INPROCESS_CRON=false on every backend process to avoid redundant polling.
PostgreSQL locks each schedule row while creating a run and advancing its
dispatch timestamp in one transaction. If scheduler processes overlap during
a restart, only one can create a run for that occurrence. The same guard covers
standalone, Celery and backend dispatchers, including schedules that allow
overlapping runs. Locks are released on commit or rollback and work through
transaction poolers.
Outside the lean Compose stack, migrating from Celery Beat to
python -m orcheo_backend.worker.cron_scheduler requires stopping the old
Beat process, disabling in-process cron on every backend, and setting
ORCHEO_CRON_DISPATCH_OWNER=scheduler on every Celery worker sharing the
database and queue. Restart those workers before starting the standalone
scheduler so already queued Beat dispatch tasks are ignored. The standalone
scheduler always publishes runs to Celery. Deployments retaining Celery Beat
should keep the default ORCHEO_CRON_DISPATCH_OWNER=celery and run no standalone
scheduler.
The backend serves precompressed Studio assets prepared after runtime settings
are substituted. Asset responses require revalidation because those substitutions
can change content without changing Vite filenames. Studio pages load on demand.
Shared volume ownership is initialized once per runtime user and directory
configuration; set ORCHEO_REPAIR_VOLUME_OWNERSHIP=true for one restart if files
were added as another user and need repair.
Deployments without browser workflows can copy
deploy/lean/docker-compose.no-browser.yml beside their Compose file and use:
This optional overlay requires Docker Compose 2.24.4 or newer. It scales both reader sidecars to zero and makes browser nodes fail with an explicit disabled message. The standard stack keeps isolated browser support enabled. Start again with the base Compose file to restore it.
The lean services use restart: unless-stopped. The backend's Docker health
check calls /api/system/ready, which tests Redis from the backend container;
workers use the lightweight orcheo.broker_healthcheck probe for broker
reachability. The scheduler uses orcheo.cron_healthcheck, which also checks its
process and a container-local heartbeat written after each successful schedule
poll. Missing or stale progress fails the probe after the greater of 30 seconds
or three dispatch intervals. Startup and shutdown clear the heartbeat; an
unexpected dispatch-loop exit stops the scheduler process so its restart policy
can recover it. The installer waits for all services to become healthy.
Redis availability is an intentional part of the lean backend container's
health status; use
/api/system/health as the backend liveness probe and /api/system/ready as
the readiness probe in an orchestrator that restarts unhealthy containers.
Docker Compose does not automatically restart a container merely because its
health check becomes unhealthy, so monitor Compose health and alert on an
unhealthy service or a stopped worker or Beat.
After a Redis or Docker network incident, recreate the lean containers if the
redis service name does not resolve from the backend:
cd ~/.orcheo/lean
docker compose up -d --no-build --force-recreate --wait --wait-timeout 120
docker compose exec backend python -m orcheo.broker_healthcheck
docker compose ps
Recreation retains the named redis_data and orcheo_data volumes. Do not use
down -v during recovery. Record docker compose version and
docker network inspect orcheo-lean_default if a service alias disappears.
Use Docker's maintained Compose plugin rather than an outdated distribution
package when diagnosing a repeat network issue.
Trigger-created runs carry a persisted dispatch flag. While the backend is up,
it checks PostgreSQL once per minute and republishes up to 20 flagged runs that
have remained pending for at least two minutes without a confirmed enqueue.
Each run is claimed across backend processes and retried no more than once
every five minutes after a failed attempt. Runs already accepted by Redis are
not republished simply because the worker queue is busy. Worker start
transitions lock the PostgreSQL row so duplicate queue messages cannot start
the same run twice. Monitor the count and age of pending runs where
dispatch_requested = TRUE; sustained growth means execution is stalled.
SELECT COUNT(*) AS pending_dispatches, MIN(created_at) AS oldest_created_at
FROM workflow_runs
WHERE status = 'pending' AND dispatch_requested = TRUE;
If Redis loses a message after acknowledging a publish, the run will still be
marked as enqueued. Automatic recovery of broker data loss after acceptance is
outside the reconciler's scope. The lean Compose service enables Redis AOF
persistence and keeps it in the redis_data volume; preserve and back up that
volume. If the broker data is lost, verify that the run is absent from the
worker queue before setting enqueue_confirmed = FALSE for that run to request
replay. This avoids creating duplicate queue messages during a normal backlog.
The concurrency quota also counts pending runs deliberately created through
the API and runs marked running. PostgreSQL-backed Celery workers now claim a
run with a unique owner token and renew its database lease every 20 seconds
from a separate thread, including while a long workflow node is running. The
lease lasts two minutes; a backend reconciler atomically fails an owned run
only after another three minutes without renewal. Lease decisions use the
PostgreSQL clock across all backend and worker processes. A worker that loses
its lease cancels its execution. If a blocking node does not stop within 30
seconds, the worker process exits so it cannot keep executing when the quota
slot is released. Its token cannot commit a late result. There is no maximum
workflow execution duration: a live worker can keep renewing its lease. If
PostgreSQL is unavailable long enough for the lease to expire, the worker
stops and recovery resumes when the database is available.
Runs already running before this ownership mechanism was deployed have no
owner token and are not automatically failed. Inspect their worker and run
history before marking them failed through POST /api/runs/{run_id}/fail.
Deliberately idle or abandoned API-created pending runs also require operator
review; mark or cancel them through the run API. To find candidates:
SELECT id, workspace_id, status, created_at, updated_at
FROM workflow_runs
WHERE status IN ('pending', 'running')
ORDER BY updated_at;
To inspect suspected worker orphans, query the lease columns:
SELECT id, workspace_id, worker_heartbeat_at, worker_lease_expires_at
FROM workflow_runs
WHERE status = 'running' AND worker_owner_token IS NOT NULL
AND worker_lease_expires_at < now();
Every five minutes the backend logs a warning for each workspace and status
with pending or running runs that have neither a state update nor a worker
heartbeat for at least one hour. The warning includes the count and oldest
update time. Alert on Stale active workflow runs in backend logs, then
inspect the worker and run history before changing run status. Long-running
runs without worker heartbeats can also trigger this warning; it does not
automatically release quota slots.
Runs created before this dispatch flag was added need operator review before replay because some API-created pending runs are deliberately idle. After checking which run IDs were meant to execute and that they have not already started, mark only those IDs for reconciliation in PostgreSQL:
UPDATE workflow_runs
SET dispatch_requested = TRUE
WHERE id IN ('reviewed-run-id-1', 'reviewed-run-id-2')
AND status = 'pending';
Cron state retains its last dispatched occurrence. After an outage, the cron
dispatcher creates at most one due occurrence per workflow per pass; schedules
with overlap protection wait for that run to finish before another is created.
If a schedule has never dispatched and has no start_at, it uses the current
time as its baseline, so occurrences from before recovery are not created.
Review the outage window for missed occurrences and the resulting backlog.
Single container¶
Without a worker or Beat, the backend runs executions and cron triggers in-process, so PostgreSQL is the only other service:
docker run -d --name orcheo -p 2025:2025 \
-v orcheo_data:/data \
-e ORCHEO_POSTGRES_DSN=postgresql://orcheo:orcheo@db.example:5432/orcheo \
-e ORCHEO_VAULT_ENCRYPTION_KEY="$(openssl rand -hex 32)" \
-e ORCHEO_AUTH_MODE=required \
-e ORCHEO_AUTH_JWT_SECRET="$(openssl rand -hex 32)" \
-e ORCHEO_STUDIO_URL=https://orcheo.example.com \
ghcr.io/ai-colleagues/orcheo-lean:latest
Studio's VITE_ORCHEO_* settings are read from the container environment at
startup, as with the stack Studio image. VITE_ORCHEO_BACKEND_URL can stay
unset because Studio calls the backend on its own origin, and
VITE_ORCHEO_APPS_BASE_DOMAIN defaults to ORCHEO_APPS_BASE_DOMAIN. Hosted
apps still need the separate app gateway, so leave ORCHEO_HOSTED_APPS_ENABLED
off with this image unless you run one. In single-container mode, run exactly
one container per database: the in-process cron loop is only safe in a single
backend process, so use the Compose setup above when you need more.
Reachable Self-Hosted Host (Bundled Caddy)¶
This is the standard public self-hosted recipe for Orcheo on a reachable Linux host. The bundled stack keeps backend, Studio, Postgres, Redis, worker, and beat on the Docker network while Caddy is the only service that needs public 80/443.
- Prepare the host
- Point your DNS hostname at the machine that will run Docker.
- Open inbound
80and443. - Install Docker and the Orcheo SDK.
- Install the stack with public ingress
- Understand the routing contract
https://orcheo.example.com/-> Studiohttps://orcheo.example.com/api/...-> backend HTTP routeswss://orcheo.example.com/ws/...-> backend WebSocket routes- Inspect the generated stack config when needed
COMPOSE_PROFILES=public-ingressenables Caddy TLS ingress. Backend and Studio remain accessible on their direct localhost ports (2025and2026by default).ORCHEO_CADDY_BACKEND_UPSTREAMScontrols the backend upstream pool for/api/*and/ws/*.- Verify the public origin
Replica Topology¶
The initial supported load-balancing topology is one logical deployment with multiple backend replicas that all share the same Postgres and Redis services. Caddy load-balances only replicas of that same deployment.
Set explicit backend upstreams in ~/.orcheo/stack/.env when you add more backend replicas:
Use this pattern only when the replicas share the same repository, checkpoint, ChatKit, and vault state through shared Postgres and Redis. Do not use one hostname and one path to multiplex isolated customer-specific stacks.
When To Put Something In Front Of Caddy¶
Bundled Caddy is appropriate for standard self-hosted installs and moderate scale. Prefer a cloud-managed load balancer, ingress controller, CDN, or WAF in front of Caddy, or instead of Caddy, when you need:
- higher-volume internet edge traffic
- managed certificates outside the host
- WAF, bot management, or DDoS shielding
- platform-native ingress on Kubernetes or managed container platforms
Published Prerelease Staging Host¶
Use the published prerelease channel when a staging host should validate the same artifacts that prerelease users will install:
The installer resolves the newest stack-vX.Y.Z-{alpha,beta,rc}.N tag, syncs
the stack assets from that exact tag, and pins both the stack and Studio images
to the resolved version. Use --stack-version instead when the host must remain
on one exact prerelease.
For unreleased source development, use the root docker-compose.yml.
Cloudflare Tunnel Or Similar Split-Origin Tunnel¶
Use this recipe when the host is not directly reachable or when you intentionally keep Studio and backend on separate public hostnames behind a tunnel. In this topology, bundled Caddy stays off and the tunnel forwards to the direct localhost ports published by backend and Studio.
- Install the stack without bundled public ingress
- Point your tunnel routes at the direct localhost ports
https://orcheo.example.com->http://localhost:2025https://orcheo-studio.example.com->http://localhost:2026- Set the generated stack env to the split-origin contract
ORCHEO_PUBLIC_INGRESS_ENABLED=false ORCHEO_API_URL=https://orcheo.example.com ORCHEO_STUDIO_URL=https://orcheo-studio.example.com VITE_ORCHEO_BACKEND_URL=https://orcheo.example.com ORCHEO_CORS_ALLOW_ORIGINS=https://orcheo-studio.example.com ORCHEO_CHATKIT_PUBLIC_BASE_URL=https://orcheo-studio.example.com VITE_ORCHEO_ALLOWED_HOSTS=localhost,127.0.0.1,orcheo-studio.example.com - Restart the stack after editing
~/.orcheo/stack/.env - Verify the public origins
The important distinction is that backend-facing values use the backend hostname, while browser-origin values use the Studio hostname. Passkeys are a browser-origin value too: they are verified against ORCHEO_STUDIO_URL, so it must be the Studio hostname. If these are collapsed back to localhost values, browsers will fail preflight requests and the backend will log OPTIONS ... 400.
Managed Hosting (PostgreSQL, async pool)¶
This deployment targets platforms such as Fly.io, Railway, or Kubernetes where Postgres is available as a managed service.
- Provision PostgreSQL
- Create a database and note the DSN, e.g.
postgresql://user:pass@host:5432/orcheo. - Ensure the
psycopg[binary,pool]andlanggraph[postgres]extras are installed (already defined inpyproject.toml). - Configure environment variables
export ORCHEO_CHECKPOINT_BACKEND=postgres export ORCHEO_POSTGRES_DSN=postgresql://user:pass@host:5432/orcheo export ORCHEO_REPOSITORY_BACKEND=postgres export ORCHEO_CHATKIT_BACKEND=postgres export ORCHEO_HOST=0.0.0.0 export ORCHEO_PORT=2025 export ORCHEO_VAULT_BACKEND=postgres export ORCHEO_VAULT_ENCRYPTION_KEY=change-me export ORCHEO_VAULT_TOKEN_TTL_SECONDS=900 - Deploy the application
- Docker image: Build with
docker build -t orcheo-app .and push to your registry. - Fly.io example:
- Ensure the container command starts uvicorn:
uvicorn orcheo_backend.app:app --host 0.0.0.0 --port ${PORT}. - Health checks
- Expose
/docsand/openapi.jsonfor HTTP checks. - Use
/ws/workflow/{workflow_id}for synthetic workflow runs during smoke tests.
Verification: Run uv run pytest tests/test_persistence.py locally with the ORCHEO_CHECKPOINT_BACKEND=postgres environment variable set and a reachable Postgres DSN to mirror production behavior.
Vault note: Managed environments should prefer KMS-integrated vaults. Configure IAM policies so only the Orcheo runtime can decrypt with the specified key.
Operational Tips¶
- Secrets: Prefer platform-specific secret managers (Fly Secrets, Railway variables, AWS Parameter Store) and never bake DSNs or vault encryption keys into images.
- Observability: Route application logs to structured logging (e.g., stdout + centralized collector) and enable OpenTelemetry tracing via the
ORCHEO_TRACING_*variables (see OpenTelemetry Tracing). - Scaling: The FastAPI app is stateless. Scale horizontally by adding replicas while pointing them at the same checkpoint database. With bundled Caddy, keep replica pools limited to one logical deployment that shares Postgres and Redis.
- Backups: Schedule database backups (pg_dump or managed snapshots) to protect workflow history and run states.
Use Cloudflare Tunnel when the host is not directly reachable from the internet, or when you intentionally want tunnel-managed public hostnames in front of the direct localhost ports. For reachable hosts with direct inbound ports and one shared origin, bundled Caddy is the simpler default.
These recipes will evolve as additional milestones introduce credential vaulting, trigger services, and observability pipelines.