Deployment
Service-level problems — startup failures, /ready never passing, migration and cleanup job failures, and connection saturation.
These are problems with the running service itself, independent of Canvas. Start
here when /health or /ready fail.
The service fails to start
In a hardened runtime (production, or any Cloud Run service), startup fails before listening if configuration is unsafe or incomplete — this is intentional.
Check the configuration:
- All required values are present: database, Canvas, LTI, OAuth, secrets, URL, and a valid encryption certificate. See Required application values.
TOOL_URLandCANVAS_DOMAINare HTTPS origins with no path/query/credentials.SESSION_SECRETandSTATE_ENCRYPTION_KEYare each ≥32 chars and different from each other.LTI_PRIVATE_KEYis a valid RSA 2048+ JWK.- You did not set both a direct secret and its
_FILEform (that is rejected). - Hardened runtimes reject
USE_IN_MEMORY_STORE=true,APP_DEBUG_ENABLED=true, and disabled encryption.
/ready never passes
/ready reflects PostgreSQL reachability and whether all checked-in migrations
are applied.
Check:
- PostgreSQL is reachable from the service (host, port, SSL mode, credentials).
- The migration job/service ran and succeeded against this exact image before traffic. A new image whose migrations have not run will fail readiness by design.
- On Cloud Run, the Cloud SQL attachment and socket path
(
/cloudsql/PROJECT:REGION:INSTANCE) are correct. - On Docker, the
migrateservice completed —appwaits on its success condition.
Migration job fails
Check:
- Checksum validation — migrations are forward-only and checksummed; if an already-applied migration was edited, validation fails. Add a new forward migration instead of editing history.
- The migration ran with the intended image and against the correct database.
- On Cloud Run, review the migration job execution logs; the service deploy waits for the migration job and will not shift traffic if it failed.
Cleanup job fails or stops running
Expired sessions, transient state, and locks accumulate if cleanup stops.
Check:
- The schedule exists and is firing (Cloud Scheduler job, or the systemd timer/cron entry).
- The scheduler identity has invoker permission on the cleanup job (Cloud Run).
- Review the cleanup log/executions. Alert on stale or failed executions.
See Operations → Scheduled cleanup.
PostgreSQL connection saturation
Check the pool math:
(max app instances × DATABASE_POOL_MAX) + job/admin reserve < database max_connectionsRecalculate before changing DATABASE_POOL_MAX or the instance count. The Cloud
Run configs use a pool max of 5; ten production instances use at most 50
application connections before the reserve. See
Pool sizing.
Verifying a deployment quickly
curl -fsS "${TOOL_URL}/health"
curl -fsS "${TOOL_URL}/ready"
curl -fsS "${TOOL_URL}/.well-known/jwks.json"
curl -fsS "${TOOL_URL}/lti/config"
curl -fsS "${TOOL_URL}/js/canvas-seb-detector.js" | headIf these pass but users still fail, the problem is in Canvas or on devices — return to the symptom index.