Version v0.4.6 of the documentation is no longer actively maintained. The site that you are currently viewing is an archived snapshot. For up-to-date documentation, see the latest version.
Helm chart
The Pro control plane installs on Kubernetes via the official Helm chart. The chart is chart-test gated, publishes multi-arch images per release, and signs them with cosign. It runs on any cluster with an external Postgres + Redis.
Install from the published OCI chart
Every release tag publishes the chart as a cosign-signed OCI artifact to
oci://ghcr.io/neochaotic/charts/leoflow (ADR 0028),
co-versioned with the release tag — so you can install a pinned version without
cloning the repo:
helm install leoflow oci://ghcr.io/neochaotic/charts/leoflow --version <x.y.z> \
-n leoflow --create-namespace \
-f values.yaml
Pass the release tag without the leading v (tag v0.4.0 → --version 0.4.0):
the chart version/appVersion move in lockstep with the tag, so this also pins
the control-plane image. Installing from a source checkout
(helm install ./helm/leoflow) is still supported for unreleased branches — see
the chart README
for both paths and the full values surface.
This operator-journey page is the entry point; the exhaustive values reference is
maintained alongside the chart source so it never drifts from values.yaml:
→ Helm chart README (including the datastore compatibility matrix).
A first-class values reference on this site is a TODO for a later migration phase.
What the chart gives you
The
/api/v2/Airflow-compatible API + UI, and the scheduler — as one process (role=all) or split intorole=api+role=scheduler(ADR 0049).A guarded HA posture:
replicaCount > 1refuses to render onto a single-writer log volume, and the PodDisruptionBudget turns itself on exactly when a second replica makes it safe — see Control-plane HA and disruption posture and the one-switchexamples/values-ha.yamlprofile.Opt-in hardening templates: HPA, NetworkPolicy, and a Prometheus ServiceMonitor.
Two things about that combination are worth knowing before you turn them on.
The metrics port has its own ingress rule, and its default allows any namespace.
networkPolicy.metricsFromis where you narrow it, and you should: the metrics listener is unauthenticated and its series carrydag_idandtask_id, so “reachable at all” is not the same bar as the mostly JWT-gated API on 8080. Point it at your Prometheus namespace. Setting it to an explicitly empty list closes the port, which is a supported choice — it was just a very poor default, because an Ingress-typed policy denies what it does not match, sonetworkPolicyplusmetrics.serviceMonitor.enabledused to give you a target that was created and never answered.The HPA needs both of its bounds set together — see Autoscaling.
TLS termination via cert-manager — see Pro TLS.
Tuning the probes
probes.liveness and probes.readiness are exposed so you can loosen them
under load: a busy scheduler can miss a 1s /healthz during a task-pod burst
and be kubelet-killed mid-run, which cascades in-flight tasks to agent_lost.
The defaults are deliberately forgiving — 5s liveness and 3s readiness timeouts,
three failures — rather than Kubernetes’ 1s/3.
One of those values has a floor the chart enforces:
$ helm upgrade --install leoflow oci://ghcr.io/neochaotic/charts/leoflow \
--set probes.readiness.timeoutSeconds=1
Error: probes.readiness.timeoutSeconds=1 is below the 3s floor. [...]
/readyz checks the database connection and asserts the schema is the one
this binary requires, and it bounds that whole check at 2s so it always answers
before the kubelet stops listening. Setting the kubelet’s timeout below that
inverts the relationship: the kubelet cancels a probe that was about to reply,
so instead of a 503 whose log line names the failing dependency, you get a
bare timeout that names nothing — precisely when you most need to know which
dependency is slow. The chart refuses the value rather than installing a cluster
that will go quiet under stress.
If you need Kubernetes to react to an unready pod faster than 3s, lower
probes.readiness.periodSeconds or failureThreshold instead. Both shorten
time-to-unready without cutting off the answer:
probes:
readiness:
timeoutSeconds: 3 # leave at or above the floor
periodSeconds: 5 # check twice as often
failureThreshold: 2 # out of rotation after two misses
The readiness check runs on its own dedicated database connection, separate from the pool serving API traffic, so a saturated control plane does not make every replica report itself unready at the same moment.
Related
- Editions & operating modes — what Pro gives you and its validation status.
- Deploy your first Pro DAG — the promotion walkthrough.
- Upgrades — upgrading a chart release safely.
- Cross-listed from the Reference section for the values surface.