Helm chart
The Pro control plane installs on Kubernetes via the official Helm chart. The chart is chart-test gated, publishes multi-arch images per release, and signs them with cosign. It runs on any cluster with an external Postgres + Redis.
Install from the published OCI chart
Every release tag publishes the chart as a cosign-signed OCI artifact to
oci://ghcr.io/neochaotic/charts/leoflow (ADR 0028),
co-versioned with the release tag — so you can install a pinned version without
cloning the repo:
helm install leoflow oci://ghcr.io/neochaotic/charts/leoflow --version <x.y.z> \
-n leoflow --create-namespace \
-f values.yaml
Pass the release tag without the leading v (tag v0.4.0 → --version 0.4.0):
the chart version/appVersion move in lockstep with the tag, so this also pins
the control-plane image. Installing from a source checkout
(helm install ./helm/leoflow) is still supported for unreleased branches — see
the chart README
for both paths and the full values surface.
This operator-journey page is the entry point; the exhaustive values reference is
maintained alongside the chart source so it never drifts from values.yaml:
→ Helm chart README (including the datastore compatibility matrix).
A first-class values reference on this site is a TODO for a later migration phase.
What the chart gives you
The
/api/v2/Airflow-compatible API + UI, and the scheduler — as one process (role=all) or split intorole=api+role=scheduler(ADR 0049).A guarded HA posture:
replicaCount > 1refuses to render onto a single-writer log volume, and the PodDisruptionBudget turns itself on exactly when a second replica makes it safe — see Control-plane HA and disruption posture and the one-switchexamples/values-ha.yamlprofile.Opt-in hardening templates: HPA, NetworkPolicy, and a Prometheus ServiceMonitor.
Two things about that combination are worth knowing before you turn them on.
The metrics port has its own ingress rule, and its default allows any namespace.
networkPolicy.metricsFromis where you narrow it, and you should: the metrics listener is unauthenticated and its series carrydag_idandtask_id, so “reachable at all” is not the same bar as the mostly JWT-gated API on 8080. Point it at your Prometheus namespace. Setting it to an explicitly empty list closes the port, which is a supported choice — it was just a very poor default, because an Ingress-typed policy denies what it does not match, sonetworkPolicyplusmetrics.serviceMonitor.enabledused to give you a target that was created and never answered.The HPA needs both of its bounds set together — see Autoscaling.
TLS termination via cert-manager — see Pro TLS.
OIDC/SSO login (
auth.oidc.*), off by default: turning it on renders the ConfigMap the two map settings need and makesauth.provider: oidcreachable from values, rather than requiring a hand-mounted config file. See SSO with Google Workspace or SSO with another OIDC provider.
Tuning the probes
The startup gate
probes.startup runs before the other two, and while it is failing Kubernetes
runs neither the liveness nor the readiness probe. That is what keeps a slow
boot from being read as an unhealthy process. The pod is still not Ready during
that window, so it stays out of the Service’s endpoints: the gate buys boot time,
it does not send traffic anywhere early.
It matters because liveness and readiness both target the API listener, and that
listener binds at the end of boot. Without a startup gate the kubelet answers
a boot that is slow or stuck by restarting the container, roughly every minute,
and each cycle is recorded as Completed exit=0 because the process handles
SIGTERM cleanly. An operator triaging that sees a Deployment whose containers
keep finishing successfully, a readiness probe refusing connections, and no cause
anywhere.
probes:
startup:
enabled: true
periodSeconds: 5
failureThreshold: 36
The budget is periodSeconds * failureThreshold, 180 seconds by default. That
number is sized from the server’s own boot bounds rather than picked: the
Postgres connect retry is 30 seconds and runs twice (the request pool and the
dedicated probe pool), the pod-informer warm-up is bounded at 10 seconds, and
OIDC discovery is bounded at 15 seconds. That is 85 seconds of code-defined
budget before anything cloud-shaped, plus headroom for what still carries no
bound of its own (credential detection for an object-store log sink, and
process start), and a gate tighter than the boot would kill a pod that was
about to come up.
It is deliberately generous, because the two directions are not symmetric. A boot
that fails returns an error and the process exits 1, so CrashLoopBackOff
reports it promptly whatever this budget says. The gate only bounds a boot that
hangs, and restarting a hang is a lottery ticket rather than a fix. The cost of
being too loose is a slower restart of something that was not going to recover
anyway; the cost of being too tight is the restart loop back.
The chart refuses to render a startup budget at or below the liveness budget it
replaces (initialDelaySeconds + periodSeconds * failureThreshold, 70 seconds by
default), because such a gate only moves the kill from one probe to the other.
Raise failureThreshold if your control plane legitimately takes longer to come
up, for example against a distant database or a node whose cloud metadata server
is not yet serving credentials. Set enabled: false to hand the boot back to
liveness deliberately.
The gate probes /healthz, not /readyz, and that is not an oversight. A
failing startup probe restarts the container, exactly like a failing liveness
probe, so gating it on dependency health would turn a database outage or an
in-flight migration into a crash loop. /healthz is a static 200 for the same
reason: restarting a pod creates no schema and reaches no database. Nothing is
lost by the narrower gate, because readiness still gates endpoint membership on
/readyz the moment the gate opens, so a pod never joins the Service on the
strength of a bound listener alone.
If a pod never passes the gate, read /readyz rather than guessing: it names the
dependency that is not ready. The common causes are a database the pod cannot
reach and a ServiceAccount that cannot list/watch pods in the task namespace.
Timeouts under load
probes.liveness and probes.readiness are exposed so you can loosen them
under load: a busy scheduler can miss a 1s /healthz during a task-pod burst
and be kubelet-killed mid-run, which cascades in-flight tasks to agent_lost.
The defaults are deliberately forgiving — 5s liveness and 3s readiness timeouts,
three failures — rather than Kubernetes’ 1s/3.
One of those values has a floor the chart enforces:
$ helm upgrade --install leoflow oci://ghcr.io/neochaotic/charts/leoflow \
--set probes.readiness.timeoutSeconds=1
Error: probes.readiness.timeoutSeconds=1 is below the 3s floor. [...]
/readyz checks the database connection and asserts the schema is the one
this binary requires, and it bounds that whole check at 2s so it always answers
before the kubelet stops listening. Setting the kubelet’s timeout below that
inverts the relationship: the kubelet cancels a probe that was about to reply,
so instead of a 503 whose log line names the failing dependency, you get a
bare timeout that names nothing — precisely when you most need to know which
dependency is slow. The chart refuses the value rather than installing a cluster
that will go quiet under stress.
If you need Kubernetes to react to an unready pod faster than 3s, lower
probes.readiness.periodSeconds or failureThreshold instead. Both shorten
time-to-unready without cutting off the answer:
probes:
readiness:
timeoutSeconds: 3 # leave at or above the floor
periodSeconds: 5 # check twice as often
failureThreshold: 2 # out of rotation after two misses
The readiness check runs on its own dedicated database connection, separate from the pool serving API traffic, so a saturated control plane does not make every replica report itself unready at the same moment.
Related
- Editions & operating modes — what Pro gives you and its validation status.
- Deploy your first Pro DAG — the promotion walkthrough.
- Upgrades — upgrading a chart release safely.
- SSO with Google Workspace or
with another OIDC provider: turning on
auth.oidc.*. - Cross-listed from the Reference section for the values surface.