EC2 power leasing across AWS accounts
A central lease service connects ServiceNow requests to EC2 automation across AWS accounts, with explicit scheduling rules and security boundaries.
Chris Baran
· 8 min read
On this page
- Centralise lease records and execute within each account
- Bootstrap deployment roles before the application
- Make scheduling precedence explicit
- Explain each security boundary on its own terms
- Separate table reads, table writes and key use
- Treat dry-run as a control over power actions
- Modernise the service without losing its history
An EC2 shutdown schedule looks simple until a team needs a machine outside normal hours, a workload follows a different time zone, or maintenance requires a stopped instance to be available. Power leasing makes those exceptions part of an explicit operating policy.
For an oil and gas company, the work involved modernising inherited automation over late 2024 and early 2025. Existing scheduling logic was carried into a shared lease service, a private ServiceNow integration and repeatable deployment across AWS accounts. The engineering challenge was to connect those parts without losing the operational rules already embedded in the application.
Centralise lease records and execute within each account
The architecture separates the request, the stored lease and the action against EC2.
In the central management account, a DynamoDB table holds the lease records and uses AWS KMS encryption. ServiceNow integrates through a private Amazon API Gateway endpoint. Lambda functions behind that API retrieve lease information and update existing records.
Each participating account contains two further Lambda functions. Discovery identifies eligible instances from their environment tags and registers them in the central table. Automation retrieves the leases for its account, examines instance state and evaluates whether a power action is required.
| Component | Responsibility |
|---|---|
| ServiceNow integration | Submit and retrieve lease information through the private API |
| Central lease table | Store the shared record of requested operating behaviour |
| Account discovery | Register eligible EC2 instances for lease management |
| Account automation | Evaluate lease conditions and issue permitted start or stop actions |
The per-account resources are distributed through self-managed CloudFormation StackSets. The template schedules discovery every minute and automation every five minutes. Deployment therefore establishes a common implementation while keeping EC2 actions within the account containing the instances.
The request path updates the lease; it does not start the instance directly. Discovery must first establish the record that the update handler expects. After an update, the scheduled automation evaluates the new intent. A successful API response confirms the record change, while the eventual power action depends on a subsequent evaluation and EC2 execution.
Bootstrap deployment roles before the application
The rollout uses two StackSets with different permission models. A service-managed StackSet installs the execution IAM role in participating accounts. A separate application StackSet then uses SELF_MANAGED permissions to deploy the account-local resources through those roles. An administration role in the deployment account completes the trust relationship.
The application’s SAM transform makes this split a practical fit. AWS does not support templates containing macros or transforms in service-managed StackSets. The execution-role bootstrap template has no transform, while the application template declares Transform: AWS::Serverless-2016-10-31. Self-managed permissions support the application’s use of that transform.
Before application rollout, SAM builds the Lambda functions and packages their artifacts into S3. The resulting template refers to uploaded artifacts, so target-account deployment can retrieve the code without access to the build machine. Packaging prepares those references; the application still uses the SAM transform.
Account selection happens during the deployment run. The pipeline lists accounts directly under each selected organisational unit (OU), then appends explicitly supplied accounts. Its bootstrap operations use an INTERSECTION filter between that compiled account list and the organisation root’s scope. The self-managed application receives the compiled list as explicit account targets.
This illustrative YAML is pseudocode for the rollout sequence. Its explanatory fields and placeholders describe the steps rather than define a CloudFormation template.
target_selection:
query: direct_accounts_in_selected_ous
append: explicit_account_additions
steps:
- prepare: build_and_package_lambda_artifacts_in_s3
- provision: administration_role_in_deployment_account
- bootstrap_execution_roles:
PermissionModel: SERVICE_MANAGED
AutoDeployment: {Enabled: true}
DeploymentTargets:
OrganizationalUnitIds: [ROOT_PLACEHOLDER]
AccountFilterType: INTERSECTION
Accounts: COMPILED_TARGETS
- deploy_application:
PermissionModel: SELF_MANAGED
Template: PACKAGED_SAM_TEMPLATE
Accounts: COMPILED_TARGETSAutomatic deployment is enabled for the IAM bootstrap. AWS handles those automatic deployments separately from account-filtered operations: account-level filters do not constrain automatic deployment to newly added accounts. That can extend IAM bootstrap coverage, but it neither rebuilds the pipeline’s explicit account list nor deploys the self-managed application. New application accounts require a pipeline rerun to refresh the selection and reconcile missing stack instances. Moving an account into an OU alone does not complete application onboarding.
For the application StackSet, CloudFormation assumes the administration role. That role has permission to assume target-account execution roles, and each execution role’s trust policy names the administration role it accepts. The execution role supplies CloudFormation’s deployment permissions. The Lambda functions run under separate runtime roles.
Deployment roles and runtime roles have different powers. The runtime automation’s production-tag condition does not constrain the deployment role. Review deployment authority and runtime access separately, including who can assume each role and what it can do.
Make scheduling precedence explicit
A lease contains more than an expiry. The implemented logic accounts for daily start and stop times, local time zones, AlwaysOn, WeekendOn and applicable patch windows.
Those conditions have an order. For instances subject to enforcement, an applicable patch window takes precedence over ordinary scheduling and can start a stopped machine for maintenance. Outside that path, the weekend check precedes AlwaysOn. A lease with weekend operation disabled can therefore lead to shutdown on a weekend even when AlwaysOn is set.
That distinction matters when translating a request form into operating behaviour. The label “always on” is easy to interpret more broadly than the code actually enforces.
This fictional YAML expresses a lease policy for discussion. Its field names are explanatory; they do not define the application’s API contract.
lease:
valid_until: "2030-08-23T18:30:00+10:00"
timezone: "Australia/Brisbane"
operating_hours:
start: "07:30"
stop: "18:30"
always_on: false
weekend_on: false
maintenance:
applicable_patch_window_takes_precedence: trueThe example makes the intended relationship visible: ordinary operation follows a local daily window, weekend operation is disabled, and a relevant maintenance window can override that schedule.
In the implementation, local time zones govern daily scheduling and weekend decisions, while patch windows are evaluated in UTC. Expired-lease cleanup resets scheduling attributes. These behaviours need to be understood together: expiry, daily hours and maintenance availability are separate inputs to the decision.
The five-minute schedule also sets the cadence of evaluation. It should not be interpreted as an exact shutdown timestamp or a guarantee that a newly requested instance will be ready within five minutes.
Explain each security boundary on its own terms
The private API provides a controlled network entry point for the ServiceNow integration. Network reachability, authority to submit a request and the AWS permissions used to execute it are separate concerns. Calling an endpoint private does not establish all three.
The central table’s resource policy uses the AWS organisation identifier as a condition on cross-account access. Role permissions and the KMS key policy also participate in that access. Together, these controls define the permitted relationships; organisation membership alone is not evidence that every permission has the narrowest possible scope.
The EC2 protection is similarly specific. The automation role’s allow for start and stop actions is conditioned on Environment != Production. That statement does not grant those actions for an instance carrying the matching Production tag value. It is a conditional allow, not a universal IAM deny. Other applicable permissions still matter when assessing effective access.
Accurate tagging is therefore part of operating the control. Discovery eligibility, the application’s scheduling decisions and the role’s permissions each perform a different job. Describing those boundaries precisely gives maintainers a clearer basis for review than a broad claim that production is protected under every possible policy combination.
Separate table reads, table writes and key use
The table’s read statement covers item reads, queries and scans for organisation principals without adding a role-name test. The write statement combines the organisation condition with role-ARN patterns for discovery and automation. Both conditions must match for that write allow to apply.
These illustrative excerpts use placeholder identifiers, fictional role patterns and abbreviated action lists. TABLE_ARN represents the central table; the surrounding policy documents are omitted.
Statement:
- Effect: Allow
Action: [dynamodb:GetItem, dynamodb:Query]
Resource: "TABLE_ARN"
Principal: "*"
Condition:
StringEquals:
aws:PrincipalOrgID: "ORG_ID_PLACEHOLDER"
- Effect: Allow
Action: [dynamodb:PutItem, dynamodb:UpdateItem, dynamodb:DeleteItem]
Resource: "TABLE_ARN"
Principal: "*"
Condition:
StringEquals:
aws:PrincipalOrgID: "ORG_ID_PLACEHOLDER"
StringLike:
aws:PrincipalArn:
- "arn:aws:iam::*:role/example-lease-discovery-*"
- "arn:aws:iam::*:role/example-lease-automation-*"Key use has its own boundary. The cross-account KMS statement combines the organisation condition with a broader application-role pattern for decryption, key description and data-key generation. That pattern is narrower than organisation membership but broader than the two table-write patterns.
Effect: Allow
Action: [kms:Decrypt, kms:DescribeKey, kms:GenerateDataKey]
Resource: "*"
Principal: "*"
Condition:
StringEquals:
aws:PrincipalOrgID: "ORG_ID_PLACEHOLDER"
StringLike:
aws:PrincipalArn: "arn:aws:iam::*:role/example-lease-*"In a KMS key policy, Resource: "*" refers to the key carrying that policy. The key’s administrative statement is separate and omitted here. Permission to use that key does not itself grant permission to write lease records; the table policy still applies to those operations.
Cross-account access depends on the caller’s identity policies as well as the table’s resource policy and the KMS key policy. The relevant permissions must align for the requested table operations and key use. Role patterns help describe which callers a policy admits, but a naming convention alone does not establish least privilege. Review the permitted actions, resources and conditions together with who can create matching roles or change their permissions.
Treat dry-run as a control over power actions
An SSM Parameter Store setting controls dry-run behaviour. When enabled, the action path logs the proposed start or stop operation and suppresses the EC2 power call.
The rest of the invocation can continue. Lease cleanup and other processing may still update or delete DynamoDB records. Dry-run suppresses EC2 power actions; it does not make the entire workflow read-only.
That scope makes it useful for inspecting scheduling decisions, provided operators treat changes to the lease table as real changes. Observing a proposed shutdown in the logs and verifying an actual EC2 transition are different checks.
A practical acceptance review would cover lease extension, expiry, local weekend boundaries, maintenance overrides and the production-tag condition. It would also distinguish the evidence gathered with power actions suppressed from the evidence needed when those actions are enabled.
Modernise the service without losing its history
Read the companion case study for the wider delivery context.
The scheduling rules were inherited. The modernisation brought that existing behaviour into a central lease model, an API for ServiceNow and a deployment structure for multiple accounts. Preserving the meaning of the existing rules was part of the engineering work.
The implemented components include discovery, scheduled enforcement, lease updates, environment-based patch-window selection and the dry-run control. Separating patch windows by operating system remained a planned enhancement. Keeping instances available during maintenance also remains distinct from installing patches: power leasing supplies availability for that process.
Future changes need the same care around rule precedence, time interpretation and the separation between recording a request and executing a power action. Those relationships give the next maintainer a concrete place to assess an extension without treating the inherited behaviour as incidental.
