
Battle-tested fixes for Terraform provisioning failures that waste hours in multi-environment setups. Covers the specific races and config traps that break fresh deploys: cloud-init timing issues, SSH connection conflicts during file transfers, Cloudflare API token format errors that only surface in staging, hardcoded domains in Caddyfiles causing cert failures, and the init-data-only-once problem with Casdoor OAuth. Every trap includes the exact error message you'll see and a copy-paste fix. Honest take: this reads like someone's incident post-mortems turned into a runbook, which makes it way more useful than generic Terraform advice. Activate it when you're debugging containers stuck in Restarting state after apply or setting up a second environment that mysteriously breaks differently than production.
npx -y skills add daymade/claude-code-skills --skill terraform-skill --agent claude-codeInstalls into .claude/skills of the current project.
Prevent a valid-looking Terraform workflow from publishing unvalidated bytes or widening a change's blast radius. Keep the user's business outcome and the actual mutation surface ahead of plan counts, green wrappers, or process completeness.
PLAN_DIGEST,
CONFIRM_*, or agent inference is not production authorization.Use these incident-derived symptom patterns to choose the next falsifying check. Confirm the current source and runtime before promoting a historical cause into the present diagnosis.
docker: not found in remote-execcloud-init still installing Docker when provisioner SSHs in.
provisioner "remote-exec" {
inline = [
"cloud-init status --wait",
"command -v docker >/dev/null || { echo 'FATAL: Docker not ready'; exit 1; }",
]
}
rsync: connection unexpectedly closed in local-execDo not infer a universal Terraform limitation from this symptom. A second SSH client can lose to the
target's connection budget, SSH policy, or a competing deploy. Keep local-exec local: package an
immutable artifact there, then use a Terraform-managed upload or a purpose-built deploy system. Give
every apply a unique remote staging path; never share /tmp/src.tar.gz across concurrent applies.
provisioner "local-exec" {
command = "tar czf /tmp/src-${self.id}.tar.gz --exclude=node_modules --exclude=.git -C ${path.module}/../../.. myproject"
}
provisioner "file" {
source = "/tmp/src-${self.id}.tar.gz"
destination = "/tmp/src-${self.id}.tar.gz"
}
provisioner "remote-exec" {
inline = ["tar xzf /tmp/src-${self.id}.tar.gz -C /data/ && rm -f /tmp/src-${self.id}.tar.gz"]
}
macOS BSD tar: --exclude must come BEFORE the source argument.
cloud-init status shows "running" foreverapt-get -y does not suppress debconf dialogs. Packages like iptables-persistent block on TTY prompts.
- |
echo iptables-persistent iptables-persistent/autosave_v4 boolean true | debconf-set-selections
echo iptables-persistent iptables-persistent/autosave_v6 boolean true | debconf-set-selections
DEBIAN_FRONTEND=noninteractive apt-get install -y iptables-persistent
Known offenders: iptables-persistent, postfix, mysql-server, wireshark-common.
EACCES: permission denied in container logs, container RestartingHost volume dirs are root-owned; container runs as non-root (uid 1001). Fix before docker compose up:
mkdir -p /data/myapp/data /data/myapp/logs
chown -R 1001:1001 /data/myapp/data /data/myapp/logs
Find UID: grep adduser.*-u or USER in Dockerfile.
Keep fail-fast behavior; attach diagnostics to failure instead of disabling set -e. Otherwise an
early failed command can be overwritten by a later green health check.
provisioner "remote-exec" {
inline = [
"set -eu",
"trap 'rc=$?; if [ $rc -ne 0 ]; then docker logs myapp --tail 20 2>&1 || true; docker ps --format \\\"table {{.Names}}\\\\t{{.Status}}\\\" || true; fi; exit $rc' EXIT",
"docker compose up -d",
"sleep 15",
"docker ps --filter name=myapp --format '{{.Status}}' | grep -q healthy || exit 1",
]
}
Restarting — database tables missingDB migrations not in provisioner. PostgreSQL docker-entrypoint-initdb.d only runs on empty data dir. Explicitly create DB + run migrations:
# After postgres healthy:
docker exec pg psql -U postgres -tc "SELECT 1 FROM pg_database WHERE datname='mydb'" | grep -q 1 \
|| docker exec pg psql -U postgres -c "CREATE DATABASE mydb;"
# Idempotent migrations:
for f in migrations/*.sql; do
VER=$(basename $f)
APPLIED=$($PSQL -tAc "SELECT 1 FROM schema_migrations WHERE version='$VER'" | tr -d ' ')
[ "$APPLIED" = "1" ] && continue
{ echo 'BEGIN;'; cat $f; echo 'COMMIT;'; } | $PSQL
$PSQL -tAc "INSERT INTO schema_migrations(version) VALUES ('$VER') ON CONFLICT DO NOTHING"
done
.envCompose interpolation gives the invoking shell higher precedence than --env-file or project .env.
An old exported value can therefore override the reviewed environment silently. Inspect what Compose
actually used; unset ambient overrides when the env file is meant to be authoritative.
# Inspect interpolation inputs and the rendered model.
docker compose --env-file .env config --environment
docker compose --env-file .env config --format json > compose.rendered.json
# Make the reviewed env file authoritative for this key.
env -u DOCKER_WITH_PROXY_MODE docker compose --env-file .env build
Invalid format for Authorization headerCaddy's Cloudflare DNS module expects a scoped API Token through Bearer authentication. Do not infer credential type, validity, or permissions from length/prefix alone. Verify the token with Cloudflare's official endpoint, then exercise the exact zone operation or provider path required by the release.
curl -fsS https://api.cloudflare.com/client/v4/user/tokens/verify \
-H "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
| jq -e '.success == true and .result.status == "active"' >/dev/null
If the credential is absent or wrong, create a least-privilege API Token through Cloudflare's current dashboard/API flow and grant only the zones/operations the provider needs. Follow the official creation contract rather than copying permission-group IDs that may drift: https://developers.cloudflare.com/fundamentals/api/get-started/create-token/.
Caddyfile or compose has literal domain names. Staging Caddy loads production config, tries to get certs for domains it doesn't own → ACME fails.
Caddyfile: Use {$VAR} — Caddy evaluates env vars at startup.
# WRONG
example.com { tls { dns cloudflare {env.CLOUDFLARE_API_TOKEN} } }
# RIGHT
{$LOBEHUB_DOMAIN} { tls { dns cloudflare {env.CLOUDFLARE_API_TOKEN} } }
Compose: Use ${VAR:?required} — fail-fast if unset or empty.
# WRONG
- APP_URL=https://example.com
# RIGHT
- APP_URL=${APP_URL:?APP_URL is required}
Pass the env var to the gateway container so Caddy can read it:
environment:
- LOBEHUB_DOMAIN=${LOBEHUB_DOMAIN:?LOBEHUB_DOMAIN is required}
- CLOUDFLARE_API_TOKEN=${CLOUDFLARE_API_TOKEN:?required for DNS-01 TLS}
Do not stop at this local assertion. Put all runtime-required keys in one schema, require the same set
from every environment file, render the exact Compose service environment, and run the exact deployed
Caddy image with that full environment before mutating live files. Caddy {$VAR} expansion can become
an empty token before parsing; a Caddyfile default is not an environment-completeness check.
Social sign in failedCasdoor init_data.json contains hardcoded redirect URIs. --createDatabase=true only applies init_data on first-ever DB creation — not on restarts. Fix via SQL in provisioner:
# Replace production domain with staging in existing Casdoor DB
$PSQL -c "UPDATE application SET redirect_uris = REPLACE(redirect_uris,
'example.com', 'staging.example.com')
WHERE name='lobechat'
AND redirect_uris LIKE '%example.com%'
AND redirect_uris NOT LIKE '%staging.example.com%';"
Also check AUTH_CASDOOR_ISSUER — it must match the Casdoor subdomain (auth.staging.example.com), not the app root domain.
Before creating a second environment, grep .tf files for hardcoded names. See references/multi-env-isolation.md for the complete matrix.
Environment isolation does not mean configuration-contract drift. Keep one required-key manifest and the same validation path for every environment. Staging and production may use different domains, credentials, instance sizes, and feature values; they must not disagree about whether a runtime key is required, optional, allowed-empty, or silently defaulted.
Will fail on apply (globally unique):
| Resource | Scope | Fix |
|---|---|---|
| SSH key pair | Region | "${env}-deploy" |
| SLS log project | Account | "${env}-logs" |
| CloudMonitor contact | Account | "${env}-ops" |
DNS duplication trap: Two environments creating A records for the same name in the same Cloudflare zone → two independent record IDs → DNS round-robin → ~50% traffic to wrong instance. Fix: use subdomain isolation (staging.example.com) or separate zones. Remember to create DNS records for ALL subdomains Caddy serves (e.g., auth.staging, minio.staging).
Snapshot cross-contamination: Unfiltered data "alicloud_ecs_snapshots" returns ALL account snapshots. New env inherits old 100GB snapshot, fails creating 40GB disk. Gate with variable:
locals {
latest_snapshot_id = var.enable_snapshot_recovery && length(local.available_snapshots) > 0
? local.available_snapshots[0].snapshot_id : null
}
Do NOT add count to the data source — changes its state address, causes drift.
Run the cheapest checks first, but do not let a syntax check certify runtime behavior. HashiCorp's
terraform validate checks syntax and internal consistency without remote state or provider APIs.
Preconditions can block before their resource action; postconditions run after change and do not undo
what already happened; check assertions warn and continue. Choose the mechanism by when damage must
be prevented.
Key checks (see references/pre-deploy-validation.md):
terraform validate.Fresh disks expose every implicit dependency. See references/zero-to-deploy-checklist.md.
Key items that break provisioners on fresh instances:
mkdir -p /data/{svc1,svc2} in cloud-init — file provisioner fails if target dir missingCREATE DATABASE — PG init scripts only run on empty data dirschema_migrations table, applied idempotentlydepends_on between resources sharing Docker networks