Skip to content
Expert Cloud & AI
Menu

SCP guardrails as code: policy layers, pipelines and tested rollout

Build AWS guardrails from reusable root and OU policies, configuration-driven deployment, protected exceptions and live validation before wider rollout.

Chris Baran

· 9 min read

On this page
  1. Separate the layers of control
  2. Keep shared responsibilities at the root
  3. Restrict a service catalogue alongside FullAWSAccess
  4. Turn OU configuration into a deployable unit
  5. Give each policy stack its own pipeline job
  6. Separate administrator restrictions from service approval
  7. Scope exceptions and protect the identities that hold them
  8. Treat public defaults and unrelated updates differently
  9. Use declarative policy for persistent service settings
  10. Validate what breaks and what still works
  11. Build the operating model alongside the policy

A useful AWS guardrail must do two things: constrain activity outside the organisation’s approved boundary and preserve the deployment paths engineers need inside it.

For an oil and gas company, this evolved from root-level security policy into a reusable OU deployment framework, workload-specific service restrictions and a new internet-exposure pilot. CloudFormation defines the controls, small configuration files select where they apply, and an Azure DevOps pipeline manages their deployment lifecycle.

This article explains how those pieces fit together. The companion case study covers the business outcome and delivery cadence. The examples below are simplified illustrations of the patterns.

Separate the layers of control

AWS Service Control Policies, or SCPs, set a ceiling on permissions for principals in member accounts. IAM policies and other applicable access controls still determine what a particular role can do. Attaching an SCP does not grant access. SCPs also do not govern the management account or permissions attached to service-linked roles. AWS explains SCP scope and effects.

That makes the placement of a control important:

LayerResponsibility
Organisation rootShared security baselines and protection of central platform responsibilities
Workload OUApproved service catalogue and additional controls for that workload category
Nominated principalsRestrictions on governance actions by specified administrator roles
Account IAMLeast-privilege permissions, trust relationships and deployment authority
Service configurationPersistent settings such as VPC, image and snapshot public-access controls
100%
Policy hierarchy showing root security controls inherited by workload OUs, additional service and administrator restrictions, account IAM permissions and a separate declarative service-configuration layer.

Root and OU policies constrain member-account permissions. Account IAM supplies access, while declarative policies govern supported service settings.

Root and OU policies constrain member-account permissions. Account IAM supplies access, while declarative policies govern supported service settings.

Keep shared responsibilities at the root

The organisation already had procurement, tagging and networking controls. The 2025 work refactored the root policies and introduced a reusable security baseline, followed by protection of deployment identities and member-account exit restrictions.

Root control areaOperating responsibility
Security and audit servicesRetain central ownership of security automation, audit collection and supporting infrastructure
Encryption and instance metadataMaintain encryption settings and constrain supported resource-creation paths
Network ownershipReserve network creation for approved platform identities
Organisation membershipKeep member accounts within the organisation’s governance boundary

CloudFormation in the management account deploys and attaches the organisation policies. StackSets provide governed execution identities across member accounts. Those are distinct parts of the system: policy deployment changes the permission ceiling, while the execution identities perform approved platform work inside the accounts.

The controls need to preserve those identities’ legitimate operations without giving ordinary workload administrators the same exceptions. The OU framework then adds workload-specific controls beneath the root baseline, using a separate deployment lifecycle.

Restrict a service catalogue alongside FullAWSAccess

A subtle SCP design choice makes the workload layer effective.

At the same node in the organisation hierarchy, an additional restrictive Allow policy does not narrow a FullAWSAccess policy that remains attached. The allows combine. An explicit deny, however, applies even when another policy allows the action. AWS documents SCP evaluation.

The workload policy therefore uses one Deny statement with NotAction. Its list combines the platform baseline and the workload’s approved actions. Anything outside that list is denied for the principals covered by the statement.

A reduced service-boundary example
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "KeepWorkloadWithinApprovedServices",
      "Effect": "Deny",
      "NotAction": [
        "sts:GetCallerIdentity",
        "cloudformation:Describe*",
        "ec2:Describe*",
        "s3:GetObject",
        "s3:PutObject"
      ],
      "Resource": "*"
    }
  ]
}

This deliberately small example shows the evaluation pattern. A working landing-zone baseline needs a substantially broader, reviewed list and the appropriate platform-automation exceptions. Resource: "*" makes this an organisation service boundary; resource-level access belongs in IAM.

Do not split the baseline and workload lists into two independent Deny/NotAction statements. Each statement would deny actions present only in the other list. Build their union into one statement instead. NotAction excludes listed actions from that statement’s effect; it does not authorise them. IAM NotAction reference.

The implemented baseline deliberately includes IAM so teams can manage workload identities inside their service boundary. A separate administrator restriction can retain central governance responsibilities. Those two controls solve different problems.

Turn OU configuration into a deployable unit

The configuration tree mirrors the AWS Organizations OU tree. Folder names identify the target path; the filename selects a CloudFormation template with the same name. One configuration/template pairing produces one policy stack.

Illustrative policy repository layout
policies/
  templates/
    service-boundary.yaml
    admin-governance.yaml
  configs/
    Workloads/
      Research/
        service-boundary.yaml
        admin-governance.yaml
      Operations/
        service-boundary.yaml

The engineer supplies workload parameters rather than hand-writing the policy or editing the pipeline:

Illustrative workload configuration
- ParameterKey: AllowedActions
  ParameterValue: >-
    "s3:GetObject","s3:PutObject","lambda:InvokeFunction"

The renderer inserts those actions alongside the fixed baseline in the template. In the actual framework, the action list is a quoted JSON fragment supplied through CloudFormation parameters. It must be validated and rendered as a complete policy before attachment.

The resolver walks the OU hierarchy through the Organizations API, matching names exactly. An explicit target-ID override supports cases where path-based resolution is unsuitable. Every pairing is validated and every target is resolved before deployment begins.

CloudFormation manages both the policy and its target attachment:

Illustrative CloudFormation policy resource
Resources:
  WorkloadBoundary:
    Type: AWS::Organizations::Policy
    Properties:
      Name: !Sub "workload-boundary-${TargetId}"
      Type: SERVICE_CONTROL_POLICY
      TargetIds:
        - !Ref TargetId
      Content: !Ref RenderedPolicyDocument

TargetId and RenderedPolicyDocument are parameters supplied to this example. The resource couples policy content with the intended attachment, making the deployment unit repeatable. CloudFormation policy resource reference.

Give each policy stack its own pipeline job

The July framework first supported folder-driven deployment, then made each stack a separate Azure DevOps job using a runtime matrix.

The pipeline has three deployment steps:

  1. Resolve: validate the full configuration tree, resolve targets and emit the matrix.
  2. Deploy: create one job per pairing, with a maximum of four concurrent jobs.
  3. Cleanup: remove owned stacks whose configuration has been deleted, after every deployment succeeds.
100%
Azure DevOps pipeline: validate templates and configuration, resolve all OU targets, deploy individual policy stacks with at most four concurrent jobs, then clean up owned orphan stacks only after successful deployment.

Validate and resolve before writes. Separate jobs expose each stack’s result; cleanup follows a successful deployment set.

Validate and resolve before writes. Separate jobs expose each stack’s result; cleanup follows a successful deployment set.

This reduced pipeline skeleton shows the dependency structure. The helper scripts and deployment credentials are supplied by the platform implementation:

Illustrative Azure DevOps deployment jobs
jobs:
  - job: Resolve
    steps:
      - bash: ./tools/emit-policy-matrix.sh
        name: Emit
 
  - job: DeployPolicy
    dependsOn: Resolve
    strategy:
      matrix: $[ dependencies.Resolve.outputs['Emit.policyMatrix'] ]
      maxParallel: 4
    steps:
      - bash: ./tools/deploy-policy.sh "$(CONFIG_REL)" "$(OU_ID)"
 
  - job: Cleanup
    dependsOn: [Resolve, DeployPolicy]
    condition: succeeded()
    steps:
      - bash: ./tools/cleanup-owned-policies.sh

A matrix leg carries the configuration path, resolved OU ID and stack name. The actual pipeline uses the platform’s AWS service connection for deployment credentials, and gates the deployment stage on successful validation and the main branch. Template, YAML and shell checks run before that stage. Azure DevOps supports this use of dependency outputs and runtime expressions.

The concurrency cap bounds deployment pressure on AWS Organizations while keeping per-stack results visible. It is not a transaction: some stacks can succeed while another fails. Cleanup waits until the full deployment set succeeds.

Ownership matters during cleanup. The framework requires both a managed stack-name prefix and its source tag before selecting a stack for deletion. It also supports a dry run, and destructive local cleanup requires an explicit operator opt-in.

Separate administrator restrictions from service approval

A service boundary controls which AWS capabilities are available. A principal-specific policy controls which governance actions particular administrators may perform.

The reusable administrator template has a fixed denied-action list and a per-OU list of principal ARN patterns. Its controls cover audit collection, log retention, model-invocation logging, creation of additional identities and selected S3 governance changes.

That permits a useful access model for an engagement: the administrator can perform approved workload tasks, while central governance actions remain restricted. The policy applies to the matching principal whether the request comes from the console, a command-line tool or a service acting on its behalf.

Keeping these as separate templates lets an OU receive the controls its operating model needs, using the same configuration and deployment lifecycle.

Scope exceptions and protect the identities that hold them

The engineering sandbox needed to create its own VPCs. Its root-policy exception was tied to a fully qualified deployment-role ARN in a specific account.

An account wildcard on an exception would let the same role name match in other accounts. Account qualification makes the exception specific to the approved platform. It still needs an appropriate trust policy and IAM permissions, and remains subject to other applicable controls.

The latest exposure pilot also protects its exempt deployment identity from workload administrators. An exception is only useful as a boundary when teams cannot rewrite the identity that holds it.

The following illustrative statement protects a deployment role, leaving its lifecycle to a separately controlled provisioner:

Illustrative protection of an exception-bearing role
Effect: Deny
Action:
  - iam:CreateRole
  - iam:DeleteRole
  - iam:UpdateAssumeRolePolicy
  - iam:AttachRolePolicy
  - iam:DetachRolePolicy
  - iam:PutRolePolicy
  - iam:DeleteRolePolicy
  - iam:PutRolePermissionsBoundary
  - iam:DeleteRolePermissionsBoundary
Resource: !Sub "arn:aws:iam::${WorkloadAccountId}:role/ApprovedNetworkDeployment"
Condition:
  ArnNotEquals:
    aws:PrincipalArn: !Ref PlatformProvisionerArn

The parameter values represent approved identities, not user-selected escape routes. Changes to their trust, permissions and policy exceptions belong to the platform’s review process. A role exception excludes that role from the specific deny; it does not override a deny inherited from another policy.

Treat public defaults and unrelated updates differently

The October pilot adds three SCP groups—network, applications and data—and one EC2 declarative policy.

The SCPs constrain exposure-related API calls: public load-balancer creation, network-edge construction, endpoint selection and supported data-service public-access settings. The declarative policy maintains service-level public-access configuration.

Condition semantics make the difference between a precise guardrail and a control that interrupts normal maintenance.

For a create operation, omitting a setting can select the service’s public default. For an update operation, omitting the setting can simply mean “do not change it.” Those cases need different treatment.

Lambda function URLs provide a compact example:

Illustrative update guard with an explicit presence check
{
  "Effect": "Deny",
  "Action": "lambda:UpdateFunctionUrlConfig",
  "Resource": "*",
  "Condition": {
    "Null": {
      "lambda:FunctionUrlAuthType": "false"
    },
    "StringNotEquals": {
      "lambda:FunctionUrlAuthType": "AWS_IAM"
    }
  }
}

The conditions are combined. An explicit change to NONE matches the deny. An update that only changes CORS and omits the auth-type key does not match this statement. Separate creation controls establish the initial authentication requirement. The policy does not itself create a private network endpoint. Lambda function URL access controls.

The same distinction matters for private REST APIs, Transfer endpoints and EKS control-plane settings. Each action must support the condition key being tested. A flag present in an API request is not automatically an IAM condition key. IAM condition-operator guidance.

Use declarative policy for persistent service settings

The pilot’s EC2 declarative policy sets VPC Block Public Access to bidirectional blocking, prevents new public AMI sharing and blocks public snapshot sharing.

Illustrative EC2 declarative VPC setting
{
  "ec2_attributes": {
    "vpc_block_public_access": {
      "internet_gateway_block": {
        "mode": { "@@assign": "block_bidirectional" },
        "exclusions_allowed": { "@@assign": "enabled" }
      }
    }
  }
}

The design allows exclusions at the service-configuration layer while constraining who can create or modify them through the network SCP. An allowed exclusion mechanism is not an exclusion that has already been created. EC2 declarative-policy syntax.

Ownership still matters in a shared VPC. Its public-access setting belongs to the account that owns the VPC, rather than the participant account running compute in shared subnets. The pilot also constrains instance launches to approved shared-subnet ownership.

VPC Block Public Access governs internet-gateway traffic; it is not a replacement for IAM, endpoint controls, firewall policy or inspection of other network paths. VPC Block Public Access guidance.

Declarative policies govern supported service configuration rather than individual principal permissions. That is particularly useful where service-linked roles are outside SCP scope. AWS explains the distinction.

Validate what breaks and what still works

Static validation is the first gate. The tooling checks action and condition-key support, renders the submitted policy and measures its size. Current quotas allow 10,240 characters per SCP and 10,000 per declarative policy, with up to ten direct attachments of each type to a target. SDK and CLI submission does not strip whitespace automatically. AWS Organizations quotas.

The next gate tests behaviour against realistic requests:

Validation layerWhat it tests
Automated policy checksExpected condition behaviour, supported keys, rendering and extraction logic
Historical replayHow proposed denies interact with the sampled deployment and service activity
Simulator cross-checksAgreement where the simulator can represent the action and context
Live denied-action checksWhether intended restrictions are effective after attachment
Private CloudFormation stacksWhether useful private deployment paths still provision
Configuration checksWhether declarative service settings are effective
MonitoringWhether new deny patterns appear as the pilot is used

The replay sampled 90 days across 62 accounts, producing 59,565 event evaluations. Requests with insufficient context remain undecidable rather than being treated as allowed. Calls made by AWS services and platform automation are considered alongside direct user activity.

The current offline suite passes 271 checks. The pilot’s live verification includes eight completed private resource stacks and a separate sweep with 52 executed checks; two additional checks were skipped. Permission-path checks and successful resource provisioning are reported separately.

100%
Validation and rollout sequence: render and test, replay historical activity, attach policies to a single-account pilot, exercise private deployments and denied operations, monitor a soak period, then consider wider rollout.

The first four stages have tooling and pilot results. The monitored soak and wider rollout are the next gates.

The first four stages have tooling and pilot results. The monitored soak and wider rollout are the next gates.

The rollout design attaches the declarative policy, then data, application and network controls in stages. It calls for a monitored soak before widening. The October work establishes the pilot, not a completed organisation-wide rollout.

Build the operating model alongside the policy

The reusable part is the combination: layered controls, a small engineer interface, controlled deployment identities, visible pipeline jobs and testing along the real delivery path.

Root baselines retain shared responsibilities. OU configuration adapts service boundaries to the workload. IAM grants access inside those limits. Declarative policy maintains supported service settings, and a pilot establishes whether the next control is ready to expand.

That is how policy becomes a platform capability engineers can keep using as the organisation adopts new AWS services and AI workloads.