ClickHouse migration job times out with custom CA bundle during upgrade

Last updated: October 6, 2026

Symptom

During a chart or version upgrade of self-hosted LangSmith with LangSmith-managed ClickHouse, the pre-upgrade ClickHouse migration job fails. The pod logs repeat Waiting for ClickHouse to be ready... and then end with:

Timeout reached. ClickHouse is not ready.

There is no further detail in the log. At the same time, the LangSmith backend is reachable and ClickHouse itself is healthy, so the failure looks contradictory: everything else works, but the migration job cannot proceed.

Cause

The migration job gates schema migrations on an authenticated HTTPS SELECT 1 check against ClickHouse, retried on an interval until it hits an overall timeout. That check runs over curl, which only trusts the CA bundle supplied through the Helm config.customCa setting.

The LangSmith backend trusts that same custom CA bundle plus the system CA store built into its image. If your custom CA bundle contains only the ClickHouse server certificate (or an outdated one) instead of the root CA that issues it, the backend still works because it falls back to the system trust store. curl in the migration job has no such fallback, so it fails certificate verification. The job then times out with the generic "not ready" message instead of a clear TLS error, because the readiness check suppresses the underlying error.

LangSmith-managed ClickHouse serves a Let's Encrypt certificate chain. The root is ISRG Root X1. The server certificate rotates about every 90 days and the intermediate changes from time to time, so a bundle pinned to either of those stops matching after a rotation. The root is valid until 2035.

This only affects deployments that set config.customCa for their managed ClickHouse connection.

Resolution

Trust ISRG Root X1, the Let's Encrypt root, rather than the ClickHouse server certificate, so the fix survives future rotations.

  1. Confirm the TLS trust issue. From the migration pod, or a backend pod with the same network access, credentials and CA mounts, run one of these. Don't use curl -v here, because it prints the request headers, including the ClickHouse password.

# Shows the TLS error, no password needed
curl -sS "https://$CLICKHOUSE_HOST:$CLICKHOUSE_PORT/ping"

# Same check as the script, password not printed
curl -sS -o /dev/null -w "%{http_code}\n" "https://$CLICKHOUSE_HOST:$CLICKHOUSE_PORT/?query=SELECT%201" \
  -H "X-ClickHouse-User: $CLICKHOUSE_USER" -H "X-ClickHouse-Key: $CLICKHOUSE_PASSWORD"

A TLS problem shows up as SSL certificate problem: unable to get local issuer certificate. Checking native-port (9440) connectivity alone won't confirm this, because the HTTP readiness check uses port 8443.

  1. Update the CA bundle. Add ISRG Root X1 to the bundle referenced by config.customCa. Don't rely on the ClickHouse server certificate alone.

  2. Check the secret. If your CA bundle is synced from an external store (such as Azure Key Vault), confirm the Kubernetes secret itself has the new certificate before rerunning the job.

  3. Rerun the upgrade or migration job.

  4. Quick check: curl -sS "https://$CLICKHOUSE_HOST:8443/ping" from the migration or backend pod should return Ok. without -k.

If you run multiple environments (staging, prod, etc.), apply the same CA bundle update to all of them before upgrading any of them, rather than fixing each one as it fails.

References