SCP guardrails as code: policy layers, pipelines and tested rollout
Build AWS guardrails from reusable root and OU policies, configuration-driven deployment, protected exceptions and live validation before wider rollout.
Chris Baran
· 9 min read
On this page
- Separate the layers of control
- Keep shared responsibilities at the root
- Restrict a service catalogue alongside FullAWSAccess
- Turn OU configuration into a deployable unit
- Give each policy stack its own pipeline job
- Separate administrator restrictions from service approval
- Scope exceptions and protect the identities that hold them
- Treat public defaults and unrelated updates differently
- Use declarative policy for persistent service settings
- Validate what breaks and what still works
- Build the operating model alongside the policy
A useful AWS guardrail must do two things: constrain activity outside the organisation’s approved boundary and preserve the deployment paths engineers need inside it.
For an oil and gas company, this evolved from root-level security policy into a reusable OU deployment framework, workload-specific service restrictions and a new internet-exposure pilot. CloudFormation defines the controls, small configuration files select where they apply, and an Azure DevOps pipeline manages their deployment lifecycle.
This article explains how those pieces fit together. The companion case study covers the business outcome and delivery cadence. The examples below are simplified illustrations of the patterns.
Separate the layers of control
AWS Service Control Policies, or SCPs, set a ceiling on permissions for principals in member accounts. IAM policies and other applicable access controls still determine what a particular role can do. Attaching an SCP does not grant access. SCPs also do not govern the management account or permissions attached to service-linked roles. AWS explains SCP scope and effects.
That makes the placement of a control important:
| Layer | Responsibility |
|---|---|
| Organisation root | Shared security baselines and protection of central platform responsibilities |
| Workload OU | Approved service catalogue and additional controls for that workload category |
| Nominated principals | Restrictions on governance actions by specified administrator roles |
| Account IAM | Least-privilege permissions, trust relationships and deployment authority |
| Service configuration | Persistent settings such as VPC, image and snapshot public-access controls |
Keep shared responsibilities at the root
The organisation already had procurement, tagging and networking controls. The 2025 work refactored the root policies and introduced a reusable security baseline, followed by protection of deployment identities and member-account exit restrictions.
| Root control area | Operating responsibility |
|---|---|
| Security and audit services | Retain central ownership of security automation, audit collection and supporting infrastructure |
| Encryption and instance metadata | Maintain encryption settings and constrain supported resource-creation paths |
| Network ownership | Reserve network creation for approved platform identities |
| Organisation membership | Keep member accounts within the organisation’s governance boundary |
CloudFormation in the management account deploys and attaches the organisation policies. StackSets provide governed execution identities across member accounts. Those are distinct parts of the system: policy deployment changes the permission ceiling, while the execution identities perform approved platform work inside the accounts.
The controls need to preserve those identities’ legitimate operations without giving ordinary workload administrators the same exceptions. The OU framework then adds workload-specific controls beneath the root baseline, using a separate deployment lifecycle.
Restrict a service catalogue alongside FullAWSAccess
A subtle SCP design choice makes the workload layer effective.
At the same node in the organisation hierarchy, an additional restrictive Allow policy does not narrow a FullAWSAccess policy that remains attached. The allows combine. An explicit deny, however, applies even when another policy allows the action. AWS documents SCP evaluation.
The workload policy therefore uses one Deny statement with NotAction. Its list combines the platform baseline and the workload’s approved actions. Anything outside that list is denied for the principals covered by the statement.
{
"Version": "2012-10-17",
"Statement": [
{
"Sid": "KeepWorkloadWithinApprovedServices",
"Effect": "Deny",
"NotAction": [
"sts:GetCallerIdentity",
"cloudformation:Describe*",
"ec2:Describe*",
"s3:GetObject",
"s3:PutObject"
],
"Resource": "*"
}
]
}This deliberately small example shows the evaluation pattern. A working landing-zone baseline needs a substantially broader, reviewed list and the appropriate platform-automation exceptions. Resource: "*" makes this an organisation service boundary; resource-level access belongs in IAM.
Do not split the baseline and workload lists into two independent Deny/NotAction statements. Each statement would deny actions present only in the other list. Build their union into one statement instead. NotAction excludes listed actions from that statement’s effect; it does not authorise them. IAM NotAction reference.
The implemented baseline deliberately includes IAM so teams can manage workload identities inside their service boundary. A separate administrator restriction can retain central governance responsibilities. Those two controls solve different problems.
Turn OU configuration into a deployable unit
The configuration tree mirrors the AWS Organizations OU tree. Folder names identify the target path; the filename selects a CloudFormation template with the same name. One configuration/template pairing produces one policy stack.
policies/
templates/
service-boundary.yaml
admin-governance.yaml
configs/
Workloads/
Research/
service-boundary.yaml
admin-governance.yaml
Operations/
service-boundary.yamlThe engineer supplies workload parameters rather than hand-writing the policy or editing the pipeline:
- ParameterKey: AllowedActions
ParameterValue: >-
"s3:GetObject","s3:PutObject","lambda:InvokeFunction"The renderer inserts those actions alongside the fixed baseline in the template. In the actual framework, the action list is a quoted JSON fragment supplied through CloudFormation parameters. It must be validated and rendered as a complete policy before attachment.
The resolver walks the OU hierarchy through the Organizations API, matching names exactly. An explicit target-ID override supports cases where path-based resolution is unsuitable. Every pairing is validated and every target is resolved before deployment begins.
CloudFormation manages both the policy and its target attachment:
Resources:
WorkloadBoundary:
Type: AWS::Organizations::Policy
Properties:
Name: !Sub "workload-boundary-${TargetId}"
Type: SERVICE_CONTROL_POLICY
TargetIds:
- !Ref TargetId
Content: !Ref RenderedPolicyDocumentTargetId and RenderedPolicyDocument are parameters supplied to this example. The resource couples policy content with the intended attachment, making the deployment unit repeatable. CloudFormation policy resource reference.
Give each policy stack its own pipeline job
The July framework first supported folder-driven deployment, then made each stack a separate Azure DevOps job using a runtime matrix.
The pipeline has three deployment steps:
- Resolve: validate the full configuration tree, resolve targets and emit the matrix.
- Deploy: create one job per pairing, with a maximum of four concurrent jobs.
- Cleanup: remove owned stacks whose configuration has been deleted, after every deployment succeeds.
This reduced pipeline skeleton shows the dependency structure. The helper scripts and deployment credentials are supplied by the platform implementation:
jobs:
- job: Resolve
steps:
- bash: ./tools/emit-policy-matrix.sh
name: Emit
- job: DeployPolicy
dependsOn: Resolve
strategy:
matrix: $[ dependencies.Resolve.outputs['Emit.policyMatrix'] ]
maxParallel: 4
steps:
- bash: ./tools/deploy-policy.sh "$(CONFIG_REL)" "$(OU_ID)"
- job: Cleanup
dependsOn: [Resolve, DeployPolicy]
condition: succeeded()
steps:
- bash: ./tools/cleanup-owned-policies.shA matrix leg carries the configuration path, resolved OU ID and stack name. The actual pipeline uses the platform’s AWS service connection for deployment credentials, and gates the deployment stage on successful validation and the main branch. Template, YAML and shell checks run before that stage. Azure DevOps supports this use of dependency outputs and runtime expressions.
The concurrency cap bounds deployment pressure on AWS Organizations while keeping per-stack results visible. It is not a transaction: some stacks can succeed while another fails. Cleanup waits until the full deployment set succeeds.
Ownership matters during cleanup. The framework requires both a managed stack-name prefix and its source tag before selecting a stack for deletion. It also supports a dry run, and destructive local cleanup requires an explicit operator opt-in.
Separate administrator restrictions from service approval
A service boundary controls which AWS capabilities are available. A principal-specific policy controls which governance actions particular administrators may perform.
The reusable administrator template has a fixed denied-action list and a per-OU list of principal ARN patterns. Its controls cover audit collection, log retention, model-invocation logging, creation of additional identities and selected S3 governance changes.
That permits a useful access model for an engagement: the administrator can perform approved workload tasks, while central governance actions remain restricted. The policy applies to the matching principal whether the request comes from the console, a command-line tool or a service acting on its behalf.
Keeping these as separate templates lets an OU receive the controls its operating model needs, using the same configuration and deployment lifecycle.
Scope exceptions and protect the identities that hold them
The engineering sandbox needed to create its own VPCs. Its root-policy exception was tied to a fully qualified deployment-role ARN in a specific account.
An account wildcard on an exception would let the same role name match in other accounts. Account qualification makes the exception specific to the approved platform. It still needs an appropriate trust policy and IAM permissions, and remains subject to other applicable controls.
The latest exposure pilot also protects its exempt deployment identity from workload administrators. An exception is only useful as a boundary when teams cannot rewrite the identity that holds it.
The following illustrative statement protects a deployment role, leaving its lifecycle to a separately controlled provisioner:
Effect: Deny
Action:
- iam:CreateRole
- iam:DeleteRole
- iam:UpdateAssumeRolePolicy
- iam:AttachRolePolicy
- iam:DetachRolePolicy
- iam:PutRolePolicy
- iam:DeleteRolePolicy
- iam:PutRolePermissionsBoundary
- iam:DeleteRolePermissionsBoundary
Resource: !Sub "arn:aws:iam::${WorkloadAccountId}:role/ApprovedNetworkDeployment"
Condition:
ArnNotEquals:
aws:PrincipalArn: !Ref PlatformProvisionerArnThe parameter values represent approved identities, not user-selected escape routes. Changes to their trust, permissions and policy exceptions belong to the platform’s review process. A role exception excludes that role from the specific deny; it does not override a deny inherited from another policy.
Treat public defaults and unrelated updates differently
The October pilot adds three SCP groups—network, applications and data—and one EC2 declarative policy.
The SCPs constrain exposure-related API calls: public load-balancer creation, network-edge construction, endpoint selection and supported data-service public-access settings. The declarative policy maintains service-level public-access configuration.
Condition semantics make the difference between a precise guardrail and a control that interrupts normal maintenance.
For a create operation, omitting a setting can select the service’s public default. For an update operation, omitting the setting can simply mean “do not change it.” Those cases need different treatment.
Lambda function URLs provide a compact example:
{
"Effect": "Deny",
"Action": "lambda:UpdateFunctionUrlConfig",
"Resource": "*",
"Condition": {
"Null": {
"lambda:FunctionUrlAuthType": "false"
},
"StringNotEquals": {
"lambda:FunctionUrlAuthType": "AWS_IAM"
}
}
}The conditions are combined. An explicit change to NONE matches the deny. An update that only changes CORS and omits the auth-type key does not match this statement. Separate creation controls establish the initial authentication requirement. The policy does not itself create a private network endpoint. Lambda function URL access controls.
The same distinction matters for private REST APIs, Transfer endpoints and EKS control-plane settings. Each action must support the condition key being tested. A flag present in an API request is not automatically an IAM condition key. IAM condition-operator guidance.
Use declarative policy for persistent service settings
The pilot’s EC2 declarative policy sets VPC Block Public Access to bidirectional blocking, prevents new public AMI sharing and blocks public snapshot sharing.
{
"ec2_attributes": {
"vpc_block_public_access": {
"internet_gateway_block": {
"mode": { "@@assign": "block_bidirectional" },
"exclusions_allowed": { "@@assign": "enabled" }
}
}
}
}The design allows exclusions at the service-configuration layer while constraining who can create or modify them through the network SCP. An allowed exclusion mechanism is not an exclusion that has already been created. EC2 declarative-policy syntax.
Ownership still matters in a shared VPC. Its public-access setting belongs to the account that owns the VPC, rather than the participant account running compute in shared subnets. The pilot also constrains instance launches to approved shared-subnet ownership.
VPC Block Public Access governs internet-gateway traffic; it is not a replacement for IAM, endpoint controls, firewall policy or inspection of other network paths. VPC Block Public Access guidance.
Declarative policies govern supported service configuration rather than individual principal permissions. That is particularly useful where service-linked roles are outside SCP scope. AWS explains the distinction.
Validate what breaks and what still works
Static validation is the first gate. The tooling checks action and condition-key support, renders the submitted policy and measures its size. Current quotas allow 10,240 characters per SCP and 10,000 per declarative policy, with up to ten direct attachments of each type to a target. SDK and CLI submission does not strip whitespace automatically. AWS Organizations quotas.
The next gate tests behaviour against realistic requests:
| Validation layer | What it tests |
|---|---|
| Automated policy checks | Expected condition behaviour, supported keys, rendering and extraction logic |
| Historical replay | How proposed denies interact with the sampled deployment and service activity |
| Simulator cross-checks | Agreement where the simulator can represent the action and context |
| Live denied-action checks | Whether intended restrictions are effective after attachment |
| Private CloudFormation stacks | Whether useful private deployment paths still provision |
| Configuration checks | Whether declarative service settings are effective |
| Monitoring | Whether new deny patterns appear as the pilot is used |
The replay sampled 90 days across 62 accounts, producing 59,565 event evaluations. Requests with insufficient context remain undecidable rather than being treated as allowed. Calls made by AWS services and platform automation are considered alongside direct user activity.
The current offline suite passes 271 checks. The pilot’s live verification includes eight completed private resource stacks and a separate sweep with 52 executed checks; two additional checks were skipped. Permission-path checks and successful resource provisioning are reported separately.
The rollout design attaches the declarative policy, then data, application and network controls in stages. It calls for a monitored soak before widening. The October work establishes the pilot, not a completed organisation-wide rollout.
Build the operating model alongside the policy
The reusable part is the combination: layered controls, a small engineer interface, controlled deployment identities, visible pipeline jobs and testing along the real delivery path.
Root baselines retain shared responsibilities. OU configuration adapts service boundaries to the workload. IAM grants access inside those limits. Declarative policy maintains supported service settings, and a pilot establishes whether the next control is ready to expand.
That is how policy becomes a platform capability engineers can keep using as the organisation adopts new AWS services and AI workloads.
