Skip to main content

Project GCP — Tier-0 Sandbox Guardrails

Project GCP — Tier-0 Sandbox Guardrails

Project GCP is the platform's Tier-0 module — the only module in this catalog applied before Services GCP. It does not deploy an application or shared application infrastructure; it hardens the project itself: which Google Cloud APIs are allowed to be enabled at all, a set of self-imposed quota ceilings, and — optionally — the creation of the sandbox project and a least-privilege deploying identity to operate inside it once created.

Deployment order:

Project GCP  →  Services GCP  →  App CloudRun / App GKE  →  Application Modules

Unlike every other module in this catalog, Project GCP is not meant to be run by the same identity, pipeline, or trainee-facing action that deploys applications — see "Two Identities, Deliberately Separate" below.


What It Manages​

FileResourceGoverns
apis.tfgoogle_project_serviceWhich Google Cloud APIs are enabled on the target project
quotas.tfgoogle_cloud_quotas_quota_preferenceSelf-imposed quota ceilings (Cloud Run regional CPU, Compute Engine regional CPU, project-wide GPUs, Redis regional memory, three Filestore tiers, a BigQuery daily-scan cap on development/production, and 81 Vertex AI accelerator quotas)
project.tfgoogle_project, google_project_iam_member, google_iam_deny_policy, google_resource_manager_lienOptional project creation, least-privilege IAM for a separate deploying identity, a deny policy blocking that identity from raising its own quota caps, and a deletion lien
validation.tfcheck blocksPlan-time regression guards (the API floor can never shrink; create_project requires its companion identity variables)
budget.tfgoogle_billing_budgetA billing budget scoped to the created project, so a spend alert names one user's sandbox

Folder-level organization policy guardrails (external-IP denial, service-account key restrictions, the gcp.restrictServiceUsage allowlist, etc.) are not managed by this module — see "Folder-Level Org Policies" below.


Why a Separate Module, Applied by a Separate Identity​

Every other module in this catalog assumes a project-scoped deploying identity. Project GCP needs organization or folder-level IAM (org-scoped orgpolicy.policyAdmin/iam.denyAdmin, serviceusage.serviceUsageAdmin-class permissions to write API allowlists and quota preferences) — a materially different trust tier than anything downstream. It is designed to be applied by a platform admin impersonating a dedicated, impersonation-only service account (rad-guardrails-admin in this platform's own deployment), never wired to a Cloud Build trigger, Cloud Function, or Pub/Sub path a trainee action could reach. An identity that can both execute arbitrary Terraform on request and rewrite the org's own guardrails would defeat the reason the guardrails exist.


Two Identities, Deliberately Separate​

IdentityVariableRole
Resource creatorresource_creator_identityImpersonated to run this module. Applies the API allowlist and quota preferences; when create_project = true, creates the project itself (GCP auto-grants it roles/owner on the project it creates — there is no documented way to suppress this).
Deploying identitydeploying_identity_emailThe identity that actually runs Services GCP / application-module deploys inside the new project once it exists. Receives the deploying_identity_bundle role set below — never Owner.
Deployed-by (human)deployed_by_emailThe human who requested the deployment. Receives a tier-dependent role bundle: read-only in sandbox and lab; read-only plus operator roles in production; read-only plus build-and-deploy roles in development.

This separation is the point: the identity that sets guardrails is never the same one — or reachable via the same user-triggered path — as the one that operates inside them.


The Additive-Floor API Allowlist​

apis.tf enables a hardcoded floor of APIs (local.baseline_required_apis in main.tf) merged with var.additional_apis — the variable can only add to the floor; nothing in this module can shrink it. validation.tf's api_floor_never_shrinks check exists purely as a regression guard against a future edit accidentally introducing a filter or subtraction.

The floor was grep-verified against Services_GCP, App_CloudRun, and App_GKE — the actual set of *.googleapis.com strings a resource in those three modules reads — not the larger list Services_GCP's own default_apis enables (which also turns on several APIs no resource in those three modules consumes, e.g. Gmail, Calendar, Docs, Drive, Vertex AI):

billingbudgets · compute · servicenetworking · sqladmin · run · cloudbuild
artifactregistry · secretmanager · iam · iamcredentials · file · redis
cloudresourcemanager · logging · monitoring · storage · cloudscheduler
certificatemanager · iap · binaryauthorization · eventarc · container
cloudkms · containeranalysis · ondemandscanning · pubsub · accesscontextmanager
clouddeploy · securitycenter · gkebackup · firestore · dns · servicedirectory
gkehub · gkeconnect · anthosconfigmanagement · anthospolicycontroller · mesh

That's 38 APIs (each entry above maps to <name>.googleapis.com), verified 2026-08-13 by counting local.baseline_required_apis in main.tf. Add anything an application-specific integration needs beyond this floor — e.g. aiplatform.googleapis.com for a Vertex AI-backed app — via additional_apis; it is automatically folded into both enablement and the allowlist output.

Two entries have a history worth knowing. alloydb was removed on 2026-08-11 when AlloyDB was dropped from the sandbox/development folder allowlists and added to the production denylist on an explicit cost decision — leaving it in the floor would fail google_project_service on any fresh project. billingbudgets was added because budget.tf's google_billing_budget call is quota-attributed to this project; without it the apply fails with a bare Error 403: The caller does not have permission, which reads as an IAM problem and is not one (confirmed live 2026-08-11 on deployments 605ab251 and d5f4b359). One removed, one added — which is why the count is still 38.

dns and servicedirectory are on the floor because GKE Autopilot cluster creation calls them unconditionally, not because any module configures them. Autopilot provisions a VPC-scoped Cloud DNS managed zone for cluster DNS, and calls Service Directory (ManagedResourceService.AddServiceBundle) during cluster bring-up. Neither is optional and neither can be opted out of. Omitting dns blocks google_container_cluster creation outright with Error 403: Request is disallowed by organization's constraints/gcp.restrictServiceUsage constraint ... attempting to use service 'dns.googleapis.com'; omitting servicedirectory leaves the cluster stuck in ERROR with the equivalent message for servicedirectory.googleapis.com. Both were confirmed live deploying a GKE module into a RAD-managed project.

This enables the API on the project. It does not by itself decide whether the API is permitted — see the next section.


Folder-Level Org Policies (Not Managed Here)​

Folder-scoped organization policy guardrails — compute.vmExternalIpAccess deny, service-account key creation/upload disabled, Cloud SQL public-IP restriction, and (the companion restriction to this module's own API allowlist) gcp.restrictServiceUsage, which controls which APIs are allowed to be enabled anywhere in the folder — were originally implemented inside this module and then removed.

Why removed: GCP allows only one policy object per (folder, constraint) pair. Every sandbox project sharing a folder had its own Terraform state believing it owned the exact same folder-scoped resource — confirmed live to cause real drift (one project's apply silently overwriting another's folder policy), and, worse, meant tearing down any one sandbox project's state would delete the shared policy for every other project in that folder.

Folder-level guardrails now live in rad-automation/scripts/02-setup-ui.sh (step 10, "Enable organization level configuration and deployments"), applied once per folder at setup time via raw gcloud org-policies set-policy calls — decoupled from any individual sandbox project's Terraform lifecycle. Every new sandbox project created under that folder inherits the policies automatically through GCP's policy hierarchy; you do not need to (and cannot, from this module) re-apply them per project.

An API enabled by this module (apis.tf) but not present in the folder's gcp.restrictServiceUsage allowlist would itself fail to enable — keep 02-setup-ui.sh's hardcoded allowlist in sync with baseline_required_apis by hand if the baseline ever changes; there is no longer an automatic computed link between the two.

compute.requireOsLogin is deliberately absent from the folder defaults until Services_GCP's NFS/Redis VM is updated to support OS Login (it currently uses IAP-SSH + key metadata instead of OS Login).


Quota Overrides​

quotas.tf uses the modern Cloud Quotas API (google_cloud_quotas_quota_preference) — the older Service Usage consumer-quota-override resource no longer exists in the current hashicorp/google provider (tofu validate rejects it outright). A quota is identified by a human-readable quota_id per service (e.g. "CpuAllocPerProjectRegion"), not a metric/unit/limit triple.

Every tier gets a set of self-imposed caps, each sized from this catalog's own proven-working ceiling rather than a guess. sandbox and lab use the base map (local.default_quota_overrides); development and production share one raised map (local.production_quota_overrides = local.development_quota_overrides).

KeyServicequota_idsandbox / labdevelopment / productionScope
cloud_run_cpu_allocationrun.googleapis.comCpuAllocPerProjectRegion16000 milli-vCPU (16 vCPU)32000all regions
compute_cpus_per_region + compute_cpus_<region>compute.googleapis.comCPUS-per-project-region24 vCPU48each of the eight permitted regions
compute_gpus_all_regionscompute.googleapis.comGPUS-ALL-REGIONS-per-project00project-wide
redis_total_memory_per_regionredis.googleapis.comTotalCapacityPerProjectPerRegion16 GB32all regions
filestore_standard_per_regionfile.googleapis.comStandardStorageGbPerRegion1024 GB1024all regions
filestore_premium_per_regionfile.googleapis.comPremiumStorageGbPerRegion2560 GB2560all regions
filestore_high_scale_ssd_per_regionfile.googleapis.comHighScaleSSDStorageGibPerRegion00all regions
bigquery_query_bytes_per_daybigquery.googleapis.comQueryUsagePerDay32768 MiB (32 GiB/day)1048576 (1 TiB/day)global
compute_disks_total_per_region / compute_disks_total_<region>compute.googleapis.comDISKS-TOTAL-GB-per-project-region4096 GB, all regions8192 GB, per permitted regionsee column
compute_hyperdisk_balanced_<region>compute.googleapis.comHDB-TOTAL-GB-per-project-region4096 GB8192per permitted region
compute_ssd_total_<region>compute.googleapis.comSSD-TOTAL-GB-per-project-region1024 GB2048per permitted region
Hyperdisk Extreme / ML / Confidential / Throughput, Local SSDcompute.googleapis.comHDX-…, HDML-…, HDB-CONFIDENTIAL-…, HDT-…, LOCAL-SSD-TOTAL-GB-per-project-region00all regions

On top of that map, every tier pins three accelerator and token families to 0, because none of them is reached by the Compute Engine GPU cap: 81 Vertex AI accelerator quotas (local.vertex_accelerator_overrides), the BigQuery ML generative token quotas (GenAiInputTokensPerDay, GenAiOutputTokensPerDay), and Cloud Run GPUs (four Nvidia…GpuAlloc… quota ids).

Points worth knowing before changing any of these:

  • The unit of cloud_run_cpu_allocation is milli-vCPU, not vCPU. A literal 16 is 0.016 vCPU and fails every real Cloud Run deploy with Quota violated: CpuAllocPerProjectRegion requested: 3000 allowed: 16. compute.googleapis.com/cpus is denominated in whole vCPU.
  • Some quotas cannot be set without a region. The Cloud Quotas API rejects an empty dimensions map on CPUS-per-project-region, SSD-TOTAL-GB-per-project-region and HDB-TOTAL-GB-per-project-region (and on DISKS-TOTAL-GB at the development value), so those caps are written once per permitted region — the eight regions the tier folders' gcp.resourceLocations policy allows. The rest use dimensions = {}, which applies them in every region.
  • Hyperdisk Balanced is held equal to the persistent-disk cap, not zeroed, because GKE Autopilot may back an ordinary standard-rwo PVC with either family.
  • dimensions is under lifecycle { ignore_changes }. A quota preference's location is fixed when it is created and the Cloud Quotas API cannot move one, so a project created under an older shape would otherwise plan a diff that can never apply.

Override an existing default's value with quota_value_overrides (keyed by the same short name, e.g. { cloud_run_cpu_allocation = 32000 } — note milli-vCPU, per the unit warning above; 32 here would be 0.032 vCPU), or add an entirely new quota cap with additional_quota_overrides. Before adding a new entry:

  1. Look up the real quota_id for the target service/project — the google_cloud_quotas_quota_infos data source, or gcloud beta quotas info list --service=<api> --project=<id>. Do not hand-type one from memory or carry over an old Service Usage metric name.
  2. Confirm whether the target value needs ignore_safety_checks set (a cap below current default/usage may require it — see the resource's own documentation for the valid enum values).
  3. Size the value deliberately below the shared org ceiling — many concurrent sandbox projects likely share one org-level quota pool, so a self-imposed cap fails fast and cheap in one sandbox instead of starving others.

Optional Project Creation​

Gated on create_project, which defaults to true — the module creates the project by default. Set it to false to bring a pre-existing project instead. When enabled:

  • google_project.tier_project creates the project under folder_id, linked to billing_account_id.

(The old google_project.sandbox address survives only as the from = side of the moved block in modules/Project_GCP/moved.tf:21; no resource of that name exists.) deletion_policy = "DELETE" so a genuinely intended tofu destroy can complete — the provider otherwise defaults to PREVENT and blocks the destroy outright even after the lien (below) has already been removed.

  • deploying_identity_bundle grants deploying_identity_email a least-privilege role set — roles/editor, roles/resourcemanager.projectIamAdmin, roles/iam.serviceAccountAdmin, roles/servicenetworking.networksAdmin, roles/pubsub.admin, roles/run.admin, roles/secretmanager.admin — assembled from five independently confirmed gaps in roles/editor alone (setIamPolicy/getIamPolicy is structurally excluded from Editor on nearly every resource type it otherwise fully manages: project IAM, Private Service Access peering, Pub/Sub topic IAM, Cloud Run service IAM, and per-secret Secret Manager IAM — the last confirmed live 2026-08-18 (#2761) and the widest of the five, since App_Common's app_iam binding fires for every application with a database, Cloud Run and GKE alike). Not yet validated against the full ~150-application-module catalog — treat it as a candidate pending a real /deploy-group-test-style campaign, not a guaranteed-sufficient bundle.

  • end_user_access_bundle grants deployed_by_email a tier-dependent role set (local.end_user_roles, project.tf). All three tiers get the read-only base (roles/browser plus logging/monitoring/run/container/cloudsql/compute.viewer and storage.objectViewer). development additionally gets a build-and-deploy set (run.developer, container.developer, cloudsql.client, storage.objectAdmin, secretmanager.secretAccessor, secretmanager.secretVersionAdder, artifactregistry.writer, monitoring.editor, errorreporting.user, cloudtrace.user), so a development end user genuinely can create and update resources; it also gets a purpose-built, permissionless rad-app-runtime service account (google_service_account.app_runtime, created only on this tier) that the user holds roles/iam.serviceAccountUser on. production is deliberately tighter than development — it adds only monitoring.editor, errorreporting.viewer and cloudtrace.user (operate, not reconfigure). No tier ever receives roles/owner, roles/cloudquotas.admin, roles/orgpolicy.policyAdmin, or project-scoped roles/iam.serviceAccountUser.

  • google_iam_deny_policy.deny_deploying_identity_quota_write blocks deploying_identity_email from raising its own quota caps back up. Quota-write cannot be removed by choosing a narrower predefined role, so a Deny Policy is the mechanism used here. Denies both the modern (cloudquotas.googleapis.com/quotas.update) and legacy (serviceusage.googleapis.com/quotas.update) write permissions for defense-in-depth.

  • google_resource_manager_lien.prevent_deletion is a real, structural block on resourcemanager.projects.delete, independent of whatever IAM roles the deploying identity holds.

validation.tf's create_project_requires_billing_and_deploying_identity check rejects create_project = true at plan time unless billing_account_id, deploying_identity_email, and deployed_by_email are all set — so a misconfigured attempt fails fast rather than partially creating a project with no deploying identity able to operate it.

Sandbox project lifecycle: fresh-create only, never recycle. This platform is not training-only — some users deploy real, non-reproducible work into these sandbox projects, so a project can hold real user data. Do not destroy-and-reuse a project ID for a new user. GCP's own soft-delete (30-day window) is retained as a safety net, but is not a fast-recycle mechanism — a restored project is unusable for up to 36 hours/3 days, so it cannot back a "return this project to a pool for the next user" flow. Real teardown means a deliberate projects.delete, not a routine tofu destroy-and-reprovision cycle.


Per-Project Billing Budget​

budget.tf creates a google_billing_budget scoped to the project this module created, so a spend alert names one user's sandbox. It is created only when all three of create_project = true, enable_project_budget = true, and a non-empty billing_account_id hold.

This is not a spend cap. A GCP budget only notifies — it never blocks an API call or detaches billing. Enforcement stays with the platform's own credit_billing_guard (15-minute poll, disables billing on arrears) and credit_project (hourly metering against the BigQuery billing export). What the budget adds is speed: both of those inherit the billing export's multi-hour latency, whereas budget thresholds fire off Google's near-real-time spend tracking, making this the fastest available signal that a specific sandbox is running away.

It exists per-project because the platform-side budget is scoped to the whole billing account and cannot be narrowed there — this module mints a fresh project ID per user at deploy time, so no static project list exists on the platform side at tofu apply time. Inside the module that creates the project, the ID is known.

Four threshold rules fire: at 50%, 90% and 100% of actual spend, plus once when Google forecasts the month will end over budget (typically days before the actual-spend rule). Credits and promotions are excluded (EXCLUDE_ALL_CREDITS) so the threshold tracks real chargeable spend. Notifications go to billing-account admins and users, deliberately not to the end user — they hold read-only IAM on their sandbox and can do nothing about an overrun, and the spend lands on the platform's account rather than theirs.

VariableDefaultDescription
enable_project_budgettrueCreates the per-project budget. Has no effect unless create_project = true and billing_account_id is set.
project_budget_amount150Monthly budget in whole currency units, sized against a typical single-user stack.
project_budget_currency"USD"Must match the billing account's own currency, or the budget is rejected.
project_budget_pubsub_topic""Optional Pub/Sub topic for programmatic notifications, as projects/<project>/topics/<topic>. Empty means email-only alerts to billing-account admins. Wiring this to an automated responder is the fastest possible reaction path.

Configuration Variables​

Group 0 (module metadata — module_description, module_dependency, credit_cost, public_access, shared_users, etc.) mirrors every other module's mandatory metadata block and is not reproduced here; see CLAUDE.md's "Group 0 metadata variables are mandatory" convention. Three of the meaningful configuration knobs live in Group 0 rather than a dedicated group, since they're closer to admin/API-surface settings than per-deployment project config:

VariableDefaultDescription
tiersandboxThe single most consequential input in this module. One of sandbox, development, production, lab (plan-time validated; lab is created only by a lab session and is never offered on a deploy form). It selects the quota map applied by quotas.tf, the end-user IAM bundle granted to deployed_by_email (local.end_user_roles in project.tf), and the folder the project is created in. development raises Cloud Run to 32000 milli-vCPU, Compute to 48 vCPU and Redis to 32 GB, adds a 1 TiB/day BigQuery scan cap, and grants a build-and-deploy role set plus a permissionless rad-app-runtime service account. production inherits development's quota map wholesale but is deliberately tighter on IAM — operate, not reconfigure. The tier does not change what a deploy costs: it sets the folder (and its org policies), the project budget alert and the purchased-credit admission floor.
additional_apis[]Additional Google Cloud APIs to enable beyond the built-in baseline. Also added to the API allowlist automatically.
resource_creator_identity""Service account to impersonate for all API calls this module makes. Leave blank to use the caller's own credentials. The caller must already hold roles/iam.serviceAccountTokenCreator on this identity.
enable_servicestruePresent for cross-module UI consistency with Services_GCP's own toggle of the same name. Deliberately not wired to anything — this module's API allowlist (apis.tf) is always enforced regardless of this value; there is no supported way to skip it.

Project Configuration​

VariableDefaultDescription
project_id(required)When create_project = true, the ID of the new project to create (6–30 chars, lowercase letters/digits/hyphens, starting with a letter); otherwise an existing project ID.
create_projecttrueCreates project_id as a new GCP project. When false, the caller brings a pre-existing project and grants deploying_identity_email access to it themselves — this module never touches that path.
billing_account_id""Billing account to link the new project to. Required when create_project = true (enforced at plan time).
deploying_identity_email""Service account that will deploy modules into the new project once it's created. Required when create_project = true. Receives the least-privilege bundle described above, never Owner.
deployed_by_email""Email of the human who requested the deployment. Required when create_project = true. Must be a real Google identity (Workspace/Gmail account) — an IAM grant to an email with no corresponding Google identity is accepted by the API but grants no real access until one exists.
folder_id"" — no defaultNumeric ID of the GCP folder that holds this tier's projects — rad-sandbox, rad-development, rad-production or rad-lab, selected by var.tier and injected by the platform — used only to place a created project (google_project.tier_project). It previously defaulted to RAD's own rad-sandbox folder id; that default was removed (#2774) because this value decides where the project is created, and a default is a wrong answer waiting to be used. Tagged {{UIMeta group=0}}, so the platform injects it rather than the operator typing it; the four tier folders deliberately carry different org policy sets.
region"us-central1"Scopes the one quota cap that cannot be applied project-wide: compute_cpus_per_region (CPUS-per-project-region) requires its region dimension, so quotas.tf sets it with { region = var.region } while every other cap uses dimensions = {} (all regions). Not user-facing (no UIMeta tag) — a platform/admin decision about where sandbox compute lands, matching every other module's region default.

(The claim "no resource reads this value" is now false: google_cloud_quotas_quota_preference.guardrails reads it through local.default_quota_overrides.compute_cpus_per_region.dimensions.) | quota_value_overrides | {} | Overrides the numeric value of a default quota override, keyed by the same name (e.g. { cloud_run_cpu_allocation = 32 }). | | additional_quota_overrides | {} | Adds quota overrides beyond the seven defaults, keyed by a short name. See "Quota Overrides" above before adding an entry. |


Outputs​

OutputDescription
enabled_apisThe full set of APIs enabled (and allowlisted) on project_id by this invocation.
quota_overrides_appliedQuota overrides applied to project_id, keyed by the same short names as default_quota_overrides / additional_quota_overrides.
created_project_idThe project ID this invocation created, or null if create_project was false (project assumed pre-existing).
deploying_identity_bundle_rolesThe exact roles granted to deploying_identity_email when create_project = true — the candidate least-privilege bundle, pending full-catalog validation.

Configuration Pitfalls & Sensible Defaults​

Risk levels: Critical (data loss, full outage, security breach) — High (service unavailable or significant degradation) — Medium (degraded function or increased cost) — Low (minor impact).

VariableSensible DefaultRiskConsequence of Incorrect Value
create_project + billing_account_id/deploying_identity_email/deployed_by_emailSet all three together, or leave create_project = falseHigh 🛡 plan-timecreate_project = true with any of the three companion variables empty is rejected at plan time (create_project_requires_billing_and_deploying_identity) rather than partially creating a project with no deploying identity able to operate it.
folder_idThe tier folder the platform injects — there is no default to fall back onCriticalA project placed in a folder that hasn't had 02-setup-ui.sh step 10 run against it inherits no folder-level org policy guardrails (external-IP denial, SA key restrictions, the API allowlist) — the project is created but effectively unguarded at the org-policy layer, and this module cannot detect or warn about that from inside a single project's state.
additional_apisAdd only what a specific integration needsMediumEvery API this module enables is also implicitly trusted to already be in the folder's gcp.restrictServiceUsage allowlist — an API added here but missing from that allowlist fails to enable outright; an API added to both becomes a standing, permanent part of the project's attack surface.
quota_value_overrides / additional_quota_overridesKeep new caps below the shared org ceilingMediumA cap set at or above the org's own shared pool provides no real guardrail — the point of a self-imposed cap is to fail fast and cheap in this sandbox before starving every other concurrent sandbox project sharing the same org-level quota pool.
deploying_identity_bundle (fixed role set, not itself a variable)—LowNot yet validated against the full ~150-application-module catalog — a deploy that needs a permission outside the seven roles in the bundle fails with a clear IAM PERMISSION_DENIED, not silently. Treat it as a candidate pending a real /deploy-group-test campaign, not a guaranteed-sufficient bundle.
Destroying and reusing a project_id for a new userNever — fresh-create onlyCriticalThis platform is not training-only; a project can hold real, non-reproducible user data. GCP's soft-delete is a 30-day safety net, not a fast-recycle mechanism (restore takes up to 36 hours/3 days) — it cannot back a "return this project to a pool" flow.

Need RAD to do something it does not do yet? Request it on the roadmap, or vote on what is already there.