Xibo on GKE Autopilot — Lab Guide
Overview
Estimated time: 45–90 minutes
Xibo is an open-source digital-signage platform whose CMS schedules and distributes layouts, playlists and media to networks of display players. This lab takes you through the full operational lifecycle of the Xibo on GKE Autopilot module on Google Cloud: deploy it, access and verify it, run it day-to-day, observe it, diagnose common problems, and tear it down.
The lab focuses on operating the GKE module and the Google Cloud platform, not on Xibo product features. For the complete list of provisioned services and every configuration input (organised by group), see the Configuration Guide — this lab deliberately does not duplicate that detail so it stays accurate over time.
Objectives
By the end of this lab you will be able to:
- Deploy the module from the RAD platform with a persistent media library, and locate the resources it provisions.
- Connect to the GKE cluster and access the running CMS.
- Perform day-2 operations — inspect, update, and manage secrets, storage and the database job.
- Observe the workload with Cloud Logging and Cloud Monitoring.
- Diagnose and resolve the most common deployment and runtime issues.
- Tear the deployment down cleanly.
Prerequisites
- Services_GCP (provides the VPC, GKE Autopilot cluster, Cloud SQL, Artifact Registry, and shared service accounts this module depends on). You do not need to deploy this yourself first — the platform automatically detects whether it already exists in the target project and provisions it before this module if not (see Task 1).
- A Google Cloud project with billing enabled.
- gcloud CLI and kubectl installed;
gcloud auth loginandgcloud auth application-default logincompleted. - Project Owner (or equivalent) IAM on the project.
- Bringing your own project? Before the first deploy into it, the deployment confirmation dialog asks you to prove you control it (Get verification code, run the commands it shows as a project Owner, then Verify) and to give the RAD deployment service account the Owner role. A project RAD creates for you needs neither.
- Advanced mode for later changes. The create form asks only for the first page of inputs (and, in a project RAD creates for you, little more than the tenant name and region). Every other input in the Configuration Guide — including the scaling and version inputs in the Day-2 tasks — is changed afterwards with Update on the deployment's page after ticking Enable advanced mode, which needs a credit balance that covers the update's estimated build cost (updates never carry a module fee). On a lab environment only an administrator can use Advanced mode.
- RAD platform access with permission to deploy modules into the project.
Set these shell variables once; every task below reuses them:
export PROJECT="<your-gcp-project-id>"
export REGION="us-central1" # the region you deploy into
Task 1 — Deploy the module [Automated]
-
Open Solutions → Solution Catalog → RAD modules in the RAD platform top navigation, open Xibo (GKE) from the Platform Modules list to start configuration, choose Configuration Form under How would you like to configure this deployment? (the form opens on the Conversational Assistant if you hold purchased credits or are a partner or administrator), set
project_id, and review the inputs. Configure only what you need — the Configuration Guide documents every input by group, with defaults. Decide on media persistence now: with the module's defaults Xibo's media library is not on a persistent volume, and a StatefulSet's volume settings cannot be changed in place later. For a deployment you intend to keep, setstateful_pvc_enabled = true,stateful_pvc_mount_path = "/var/www/cms/library", astateful_pvc_sizethat fits your media, andmax_instance_count = 1(these may only be reachable in Advanced mode; see Prerequisites). Click Deploy Module, review the estimated cost in the Deployment Confirmation dialog when it appears and click Submit (if the dialog then adds a confirmation step, such as verifying a project you bring, complete it and click Confirm), which opens the deployment status page with real-time logs. -
The platform builds the Xibo image (a thin wrapper over
ghcr.io/xibosignage/xibo-cms) with Cloud Build, provisions a Cloud SQL (MySQL 8.0) database with its Secret Manager secrets, runs the one-shotdb-initjob, and deploys the CMS into the GKE Autopilot cluster behind a Gateway. On first boot Xibo installs its own schema. First deploys take roughly 20–35 minutes (Cloud SQL creation dominates). -
Connect to the cluster and discover the namespace with name-agnostic filters:
CLUSTER=$(gcloud container clusters list --project="$PROJECT" --format="value(name)" --limit=1)
gcloud container clusters get-credentials "$CLUSTER" --region="$REGION" --project="$PROJECT"
NS=$(kubectl get ns -o name | grep xibo | head -1 | cut -d/ -f2)
echo "Cluster: $CLUSTER Namespace: $NS"
kubectl get all -n "$NS"
Task 2 — Access & verify [Manual]
-
Confirm the workload is running and the database job completed:
kubectl get deploy,statefulset,pods,pvc,jobs -n "$NS"
POD=$(kubectl get pods -n "$NS" -o name | grep -v db-init | head -1 | cut -d/ -f2)
kubectl logs -n "$NS" "$POD" | grep "\[startup\]"The
[startup]line shows the database host and port (the Cloud SQL private IP on3306) and theserver_namederived for players. -
Check whether the media library is on a persistent volume:
kubectl exec -n "$NS" "$POD" -- df -h /var/www/cms/libraryIf this shows the container's overlay filesystem rather than a mounted volume, uploads will be lost when the pod is replaced — see Task 5.
-
Find the URL. On the deployment's page in the RAD platform, read the
service_urloutput; it is the Gateway'snip.ioURL unless you set a custom domain. Check the login page:SERVICE_URL="<service_url from the deployment outputs>"
curl -s -o /dev/null -w "%{http_code}\n" "$SERVICE_URL/login" # expect 200 -
Open the CMS over HTTPS in a browser (session cookies are marked
Secure, so a login over plain HTTP does not stick). The image seeds a fixedxibo_adminaccount; the module does not set its password. Sign in with the initial credentials Xibo documents for its Docker image and change the password immediately. The database password used by the CMS is in Secret Manager:gcloud secrets list --project="$PROJECT" --filter="name~xibo" --format="value(name)"
Task 3 — Operate & keep it running (Day-2) [Manual]
-
Inspect the workload — the CMS workload, its pod and (if enabled) its persistent volume:
kubectl get deploy,statefulset,pods,pvc -n "$NS"
kubectl describe pod -n "$NS" "$POD" -
Keep one replica. Each replica would have its own library, so leave
min_instance_count/max_instance_countat1. Changes go through Update on the deployment details page — the module owns the workload spec, so a manualkubectl scalewould be reverted on the next apply. -
Update the Xibo release by changing
application_versionvia Update to another exact tag published onghcr.io/xibosignage/xibo-cms. A new image builds, the pod is replaced, and Xibo's entrypoint migrates the schema on boot. -
Manage secrets, storage, and jobs:
kubectl get secrets -n "$NS"
gcloud secrets list --project="$PROJECT" --filter="name~xibo"
kubectl get jobs -n "$NS" # db-init and any scheduled jobs -
Open a database session for inspection or maintenance:
INSTANCE=$(gcloud sql instances list --project="$PROJECT" --format="value(name)" --limit=1)
DB_USER=$(gcloud sql users list --instance="$INSTANCE" --project="$PROJECT" \
--format="value(name)" --filter="name~^xibo" --limit=1)
gcloud sql connect "$INSTANCE" --user="$DB_USER" --project="$PROJECT" -
Custom domain. If you serve the CMS on your own domain, also set
CMS_SERVER_NAMEto that host inenvironment_variables— Xibo writes it into the configuration players receive.
Task 4 — Observe: Logging & Monitoring [Manual]
-
Logs — from
kubectlor the Logs Explorer:kubectl logs -n "$NS" "$POD" --tail=50Logs Explorer filter:
resource.type="k8s_container" AND resource.labels.namespace_name="<namespace>". -
Monitoring — open the GKE / Kubernetes dashboards and review pod CPU and memory utilisation, restart counts, and request metrics. The module also provisions an uptime check (when enabled); review Monitoring → Uptime checks and Alerting → Policies.
Task 5 — Troubleshoot & debug [Manual]
Durable techniques for the failure modes you are most likely to hit. These are platform-level diagnostics and do not change with Xibo releases.
- Pod not Ready / CrashLoopBackOff: inspect events and logs. The probes target
/login; a probe path of/healthzreturns 404 and restarts a healthy pod.kubectl describe pod -n "$NS" <pod> # Events section shows scheduling/probe/mount errors
kubectl logs -n "$NS" <pod> --previous # logs from the crashed container - Database connection errors: confirm the Cloud SQL instance is
RUNNABLE, thedb-initjob completed, and the[startup]line shows the private IP on3306. IfMYSQL_ATTR_SSL_VERIFY_SERVER_CERTwas overridden totrue, PDO refuses the Cloud SQL connection. - Initialisation job failed: inspect the job and its pod logs:
kubectl get jobs -n "$NS"
kubectl logs -n "$NS" job/<job-name> - Uploaded media disappeared after a restart: the library was not on a
persistent volume (Task 2, step 2). The fix is
stateful_pvc_enabled = truewithstateful_pvc_mount_path = "/var/www/cms/library"; enabling it on an existing deployment replaces the workload with a StatefulSet whose library starts empty, so re-upload media afterwards. - Login does not stick: use the HTTPS URL; cookies are
Secure. - Players point at the wrong address: check
server_namein the[startup]line and setCMS_SERVER_NAMEfor a custom domain. - Image build failed: review Cloud Build → History; confirm the
application_versiontag exists onghcr.io/xibosignage/xibo-cms(not Docker Hub).
See the Configuration Guide's Configuration Pitfalls section for setting-specific gotchas.
Task 6 — Tear down [Automated]
On the Deployments page, open the deployment and click the Trash icon (Delete). Delete runs terraform destroy and is irreversible (the deployment record is retained for history). If a deployment is stuck and the RAD platform can no longer manage it (for example after manual changes that conflict with the Terraform state), use Purge instead (from the same Delete dialog) — it removes the deployment from RAD's records without destroying the cloud resources (it makes RAD forget the deployment). Delete removes everything the module created — the Kubernetes workload
and namespace (including any library volume), Cloud SQL database, Secret Manager
secrets, GCS buckets, and Artifact Registry images. Resources owned by
Services_GCP (the VPC, GKE cluster, shared Cloud SQL, registry) are managed
separately and are not removed here.
Summary
| Task | Type | Outcome |
|---|---|---|
| 1 — Deploy | Automated | Module builds the Xibo image, provisions Cloud SQL (MySQL 8.0) and secrets, runs db-init, and deploys the CMS |
| 2 — Access & verify | Manual | Connect to the cluster; library persistence checked; /login returns 200; admin password changed |
| 3 — Operate | Manual | Inspect workload, keep one replica, update the release, manage secrets/storage/jobs, DB access |
| 4 — Observe | Manual | Query Cloud Logging; review Cloud Monitoring metrics and uptime check |
| 5 — Troubleshoot | Manual | Diagnose pod, database, init-job, lost-media, login, player-address and build issues |
| 6 — Tear down | Automated | Delete (Trash) removes all module resources |
Need RAD to do something it does not do yet? Request it on the roadmap, or vote on what is already there.