Unified Platform Policy Framework
A GitOps policy operating model that governs Terraform Enterprise, Google Cloud, and GKE with native policy engines, progressive enforcement, actionable PR feedback, and auditable exceptions.
Problem
Cloud and Kubernetes guardrails were enforced by different engines at different lifecycle stages. Terraform Enterprise needed plan-time controls, Google Cloud needed restrictions that applied across the resource hierarchy, and GKE needed admission and audit policies after the removal of PodSecurityPolicy.
Treating these as independent policy collections created inconsistent ownership, rollout, testing, exceptions, notifications, and troubleshooting. Attempting to hide them behind one generic policy language would make simple rules look consistent while obscuring the behavior that matters when a deployment is blocked.
What I built
I built a policy framework that sits above three native enforcement layers without replacing them. The framework standardizes how policies are proposed, tested, reviewed, deployed, promoted from observation to enforcement, explained to consumers, and granted exceptions.
Policy authors continue to use the language of the platform they govern: Sentinel for Terraform Enterprise, Rego and Gatekeeper resources for GKE, and native Google Cloud Organization Policy definitions. A shared GitOps operating model provides consistency while preserving engine-specific capabilities and troubleshooting paths.
Policy architecture
The framework standardizes the policy lifecycle and consumer experience while each platform remains responsible for native enforcement.
Three layers of defense
Terraform Enterprise and Sentinel
Sentinel evaluates proposed Terraform plans before cloud resources are changed. Global policy sets hold security and reliability controls owned jointly by the platform and security teams. Other platform and application teams maintain separately scoped policy sets for the resources they own.
Changes to global policy directories require independent platform and security review. Team-specific directories use their owning team's review path. After merge, repository automation updates the applicable Sentinel policy-set resources in Terraform Enterprise so the next affected plan uses the approved policy version.
Google Cloud Organization Policy
Organization Policy provides the final Google Cloud service-level boundary for centrally owned restrictions. These controls apply through the organization, folder, and project hierarchy and are managed only by the platform team in partnership with security.
A merged policy change starts a dedicated Terraform Enterprise workspace run that deploys the approved Organization Policy configuration. Application and partner platform teams cannot modify this layer directly.
GKE and OPA Gatekeeper
Gatekeeper replaced the organization's PodSecurityPolicy controls during the GKE migration. Constraint templates and constraints enforce Kubernetes-specific requirements through admission and audit. Flux reconciles the approved policy repository into clusters using the same GitOps model used for other day-one platform applications.
Gatekeeper dry-run and warning modes allow a constraint to observe existing and proposed workloads before admission changes to deny. Audit results also identify existing resources that violate newly introduced policy.
Progressive enforcement lifecycle
Every control moves through a deliberate lifecycle rather than becoming blocking immediately:
- The author defines the owner, scope, intent, native policy, remediation guidance, enforcement target, and exception behavior.
- Positive and negative unit-test fixtures verify compliant, non-compliant, boundary, and exception cases.
- Pull-request automation validates syntax, tests, repository placement, and required ownership.
- The policy is deployed in advisory, warning, or dry-run mode to measure affected resources and identify false positives.
- Affected teams receive evidence and remediation guidance before enforcement.
- Policy owners review the observation period, known impact, remediation readiness, and approved exceptions.
- The policy is promoted to blocking enforcement through another reviewed Git change.
- Deployment and evaluation status are reconciled so a failed rollout cannot appear successfully enforced.
The lifecycle is consistent, but its native mapping differs: Sentinel uses advisory, soft-mandatory, or hard-mandatory behavior; Gatekeeper uses dry-run, warn, admission deny, and audit; Organization Policy uses dry-run and active enforcement where supported.
Why Sentinel for Terraform Enterprise
The framework intentionally does not standardize Terraform controls on OPA merely because GKE already uses Rego. A Gatekeeper policy evaluates Kubernetes admission objects, while a Terraform policy evaluates a proposed infrastructure plan; sharing a language would not make their input models or policy libraries portable.
Sentinel integrates with Terraform plan, configuration, state, run, and cost context. Its advisory, soft-mandatory, and hard-mandatory model matched the required progression from observation to an auditable override and finally to controls that cannot be bypassed through the normal run workflow.
Terraform Enterprise also supports OPA policy evaluation, but its advisory and mandatory model does not preserve the same policy-level distinction between soft-mandatory and hard-mandatory checks. Choosing Sentinel retained the enforcement semantics, run visibility, and troubleshooting experience already native to the Terraform workflow.
Why native policy languages
I deliberately avoided a YAML abstraction that described a resource, attribute, and required value and then generated Sentinel, Rego, or Organization Policy code. That approach works for simple attribute checks but quickly becomes a new policy language and compiler owned by the platform team.
A lowest-common-denominator schema would lose or delay access to:
- Terraform plan, state, run, cost, and aggregate multi-resource logic.
- Gatekeeper admission matching, constraint templates, audit, and Kubernetes-specific relationships.
- Google Cloud hierarchy inheritance, conditions, tags, managed constraints, and service-defined behavior.
- Native enforcement modes, testing tools, diagnostics, and newly released engine capabilities.
The abstraction would also introduce schema versioning, three target generators, semantic-parity tests, source maps, migration logic, compatibility releases, and escape hatches. Consumers would eventually debug both the YAML representation and the generated native policy.
Instead, the framework standardizes repository structure, metadata, ownership, tests, rollout, deployment, exceptions, notifications, and troubleshooting. Cursor and Claude repository skills help contributors understand where a policy belongs and how to author it, but the native code, unit tests, and policy engine remain authoritative.
PR-level troubleshooting
A Python webhook service listens for policy failures associated with infrastructure pull requests. It correlates the exact policy version, evaluation logs, affected resource, and proposed code, then posts an AI-assisted explanation and suggested correction on the PR.
The bot comments only. It cannot push a commit, approve a change, merge code, or grant an exception. The consumer retains full control to implement the suggestion, choose another compliant design, or request an exception.
AI improves the explanation and maps the failure to the proposed code, while native evaluation and policy tests remain the source of truth.
Exception governance
Consumers can request an exception through the Platform Console or Slack, but both entry points feed the same governed workflow. Security decides whether the exception is appropriate in partnership with the owning policy and platform teams.
- Requests include the failed control, resource scope, owner, business justification, duration, and proposed compensating controls.
- Temporary exceptions expire automatically and notify the affected team and policy owners before removal.
- Permanent exceptions have no automatic expiry but are reviewed quarterly by security and the policy-owning team.
- Every policy evaluation checks the applicable exception state and records when an exception is used.
- Notifications go to both the affected application team and the policy owner, with repeated failures grouped to avoid alert fatigue.
- An exception is implemented only in the enforcement layer that owns the blocking control and cannot silently bypass a stronger layer.
Ownership and change control
Global policies are additive and cannot be weakened by a team-specific policy set. Application and partner platform teams can add stricter controls for resources they own, but cannot change organization-wide restrictions or exclude themselves from global security and reliability policy sets.
GitHub reviews provide the separation-of-duties workflow. Because standard CODEOWNERS behavior can be satisfied by one of several listed owners, the durable implementation should enforce platform and security approval independently through team-specific repository rulesets or a required approval status check for global directories.
Design tradeoffs
Using three native policy engines requires engineers to maintain Sentinel, Rego, and Organization Policy expertise. That cost is intentional: policy authors and operators can inspect the exact control executed by the target platform, use native tooling, and troubleshoot without translating through a custom abstraction.
Layered enforcement can produce overlapping failures. The shared policy metadata therefore identifies the control owner, primary enforcement layer, secondary defense, remediation path, and exception semantics so consumers receive one coherent explanation.
Progressive rollout takes longer than immediately enabling a blocking policy, but it exposes false positives, gives teams time to remediate, and makes enforcement a managed product change instead of a surprise outage.
Impact
The framework turned three policy technologies into one operating model for secure platform delivery. Platform consumers receive guardrails before deployment, actionable feedback at the pull request, and a consistent exception path. Policy owners gain versioned code, tested changes, controlled promotion, GitOps delivery, evaluation evidence, and visibility into every exception.
This strengthened defense in depth across Terraform-managed Google Cloud resources, organization-level cloud restrictions, and GKE workloads without hiding the native engines teams depend on when a policy blocks delivery.