Skip to content
Expert Cloud & AI
Menu

An enterprise AI application needs a delivery system

Principal engineering for an AI application’s delivery pipeline: local feedback, blocking checks, policy tests and usable security results, with release acceptance kept explicit.

Client: An oil and gas company

Starting expertise
Principal engineering across delivery pipelines, AWS infrastructure as code and application security.
Delivery
6–8 May 2026: one principal engineer integrated developer tooling, pipeline controls and security-result reporting for an existing AI application.
What the AI did
AI was the application workload being enabled. The contribution covered release, developer and security engineering.
What a human verified
Engineering review of gate behaviour, policy rules, negative fixtures and the presentation of scanner results.
Controls
Blocking checks, policy as code, deliberate failure fixtures, reviewable scan reports and explicit prerequisites for signing and AI evaluation.
Outcome
A versioned delivery-control foundation for an existing AI project; production release and application acceptance remained separate milestones.

An enterprise AI application needs a dependable way to accept changes. Developers need useful feedback. Security reviewers need checks with clear failure conditions. The person approving a release needs to understand what was tested, which controls ran and what still requires a decision.

Between 6 and 8 May 2026, one principal engineer implemented local developer checks, pipeline gates and security-result integration around an existing AI application for an oil and gas company. The work combined Azure Pipelines with AWS infrastructure checks. The application’s runtime remained the responsibility of its wider engineering team.

The result was a versioned control foundation for continued development. The delivery sequence made the scope concrete:

DateFocusResult
6 MayEstablish developer feedback and policyWritten design, local check entry points, infrastructure policy tests and deliberate failure fixtures.
7 MayIntegrate pipeline controls and resultsBuild, Quality, Security and conditional Sign stages, with scanner reporting and code-scanning presentation.
8 MayConsolidate the delivery foundationThe pipeline hardening and reporting changes brought together as one maintained implementation.

This window covers the developer tooling and control integration. Signing, AI evaluation and production acceptance each retained their own prerequisites.

Give developers feedback before a release is at stake

A blocking pipeline is easier to work with when developers can exercise relevant checks before submitting a change. Otherwise, routine development becomes a cycle of pushing code, waiting for a remote job and searching its logs.

The contribution added local entry points for linting, tests and security checks, together with pre-commit hooks. A lighter path supported frequent development checks; broader checks included infrastructure policy and dependency information. The central pipeline remained the acceptance boundary, since local hooks can be bypassed and workstation environments differ.

This division matters commercially. Fast local feedback helps an engineer act while the change is still fresh. Central enforcement gives reviewers a shared result. Both need maintenance: tool versions, rule configuration, dependencies and failure thresholds must stay understandable as the application grows.

Pinned tool versions make changes to the checking machinery reviewable. Vulnerability databases and downloaded rules can still change independently, so repeatability also requires knowing which inputs a scan used. A green result applies to that analysis, not indefinitely to the software.

Give each stage a clear responsibility

The pipeline was organised into Build, Quality, Security and Sign. Build prepared the available application components and synthesised AWS infrastructure templates. Quality and Security depended on Build and could run alongside one another. Sign depended on the preceding stages and on its identity configuration being ready.

That structure separated several decisions: whether the software builds, whether its tests and code checks pass, whether it meets the configured security rules, and whether an artifact can be signed. Blocking failures replaced advisory behaviour for configured checks. Scaffold-aware jobs accommodated components that were not yet present.

The diagram below is a fictional reference view of those responsibilities. It also shows the separate release decision that an organisation must establish before deployment.

100%
A fictional delivery flow runs local checks before Build, then branches into Quality and Security. Security results return to engineers through a findings view. Required checks lead to conditional signing; a separate boundary contains release approval and deployment validation.

Checks, readable findings and artifact signing support a release decision. Approval and validation of the deployed application remain separate responsibilities.

Checks, readable findings and artifact signing support a release decision. Approval and validation of the deployed application remain separate responsibilities.

A stage’s name does not establish its coverage. If a component has no tests yet, or an integration is waiting for credentials, the acceptance decision must account for that gap. This is particularly important for AI: ordinary software checks do not establish the quality of generated answers.

Keep checking after the first rejection

Independent checks should produce useful results in the same run. If one scanner rejects a change, a second scanner can still identify a different issue. Reporting must also survive a rejected build, because that is when developers most need the explanation.

The integration work made a second analysis pass eligible to run after an earlier scanner failed, and kept result publication separate from the blocking decision. Azure Pipelines normally skips later steps after failure; succeededOrFailed() changes that behaviour without requiring a successful earlier step, while still respecting cancellation. See Microsoft’s pipeline condition semantics.

This illustrative fragment uses fictional scanner wrappers. Each writes SARIF into the named results directory and returns a non-zero status for a blocking result:

steps:
  - script: ./ci/check-source.sh
    displayName: Check source
  - script: ./ci/check-infrastructure.sh
    displayName: Check infrastructure
    condition: succeededOrFailed()
  - task: AdvancedSecurity-Publish@1
    inputs:
      SarifsInputDirectory: '$(Build.ArtifactStagingDirectory)/scan-results'
    condition: succeededOrFailed()

Tool setup and report normalisation are omitted here. The important contract is that continued analysis does not forgive a failed gate. The publication task sends reports to the findings service; it does not replace the scripts’ pass-or-fail decisions. Microsoft documents the task’s inputs in the Advanced Security publishing reference.

Make the result useful to the person fixing it

A downloadable scan report is valuable, but developers also need a finding they can read in their normal review interface. The work connected scanner output to Azure DevOps code scanning through SARIF, the standard format for exchanging static-analysis results.

A shared post-processing step converted embedded HTML descriptions into Markdown for display. It preserved the original text and handled tools whose descriptions were already usable. That kept presentation work in one place instead of scattering scanner-specific fixes throughout the pipeline.

This is an engineering responsibility with a direct effect on the workflow. A reviewer needs the rule, source location, explanation and relevant reference to decide what to change. Formatting should clarify that information without changing the finding’s meaning or concealing its severity.

Microsoft’s third-party code-scanning integration guide describes SARIF ingestion and how report properties populate alerts. A readable alert is a useful handoff to an engineer; resolution still requires an engineering decision.

Treat policy and the checking machinery as software

General scanners cover common patterns. An organisation also needs executable rules for its own infrastructure expectations. The contribution added Conftest policies and policy tests against generated infrastructure configuration, bringing those requirements into code review.

For example, a fictional document service might require explicit S3 public-access blocking in every bucket template. AWS exposes four separate properties in CloudFormation’s public-access-block configuration. A policy can require all four to be enabled and reject missing configuration. That checks the proposed template; deployed permissions and account-level controls still need their own validation.

Policy tests need both allowed and rejected examples. A compliant fixture demonstrates the intended path; removing a required property demonstrates the boundary. Conftest’s policy-testing documentation explains how to verify the rules themselves as well as evaluate configuration against them.

The work also added a separate self-test workflow and deliberately invalid fixtures to exercise scanner failure paths. These fixtures belong in a controlled test area, with intentional exceptions kept separate from the application’s normal scanning scope.

When adopting this pattern, make the expected rejection precise. A scanner that cannot load its rules has failed to perform the check. A completed scan that detects the intended violation has exercised the control. This fictional Python assertion illustrates a wrapper contract that reports execution status, assigns exit code 1 to policy rejection and identifies the detected rules:

def assert_expected_rejection(report, expected_rule):
    assert report["execution"] == "completed"
    assert report["exit_code"] == 1
    assert expected_rule in report["rule_ids"]
 
assert_expected_rejection(
    {
        "execution": "completed",
        "exit_code": 1,
        "rule_ids": ["STORAGE_PUBLIC_ACCESS"],
    },
    "STORAGE_PUBLIC_ACCESS",
)

The same discipline applies to exceptions. A waiver needs a defined scope, reason, review owner and expiry. Replacing the baseline or suppressing a rule changes the acceptance boundary and deserves review alongside the application change.

Keep signing, approval and operation distinct

Artifact signing was included as a conditional stage. Activating it required the appropriate federated identity setup. Its presence in the pipeline did not mean every artifact had been signed or accepted for release.

Signature verification establishes a relationship between an artifact and a trusted signing identity. The verifier must check the expected identity and issuer; those values depend on the configured provider and trust policy. Sigstore’s signature verification guidance describes those checks. A signature alone does not establish application correctness, security completeness or permission to deploy.

Federation also needs configuration-specific claims. Microsoft’s service-connection guidance requires issuer, subject and audience values to match the configured identity. A signing integration must separately establish that its issuer and token are accepted by the signing service.

Release approval is another control. Azure Pipelines supports approvals and checks managed by resource owners, outside the pipeline YAML. A production delivery design must decide who controls those resources, which checks are mandatory and how the accepted artifact reaches the environment.

AI behaviour adds a further acceptance obligation. The pipeline reserved an evaluation job, conditional on its AWS service connection and relevant tests being available. Meaningful acceptance requires representative inputs, defined evaluation criteria and human review appropriate to the use case. A reserved job or smoke check cannot establish factual accuracy or suitability for a business decision.

Turn readiness into specific decisions

This contribution gave an existing AI project a more explicit delivery contract: local feedback, blocking checks, infrastructure policies, deliberate failure fixtures and a usable route from scan results to engineering action. It established the checking machinery while leaving application rollout, operating ownership and production acceptance as distinct work.

For a sponsor, the useful questions become concrete. Which checks ran on this candidate? Which prerequisites remain open? Who can accept an exception? What must the application demonstrate before people depend on its output?

Those questions connect this release-engineering work to AI-assisted application modernisation and bounded AI agents with verifiable results. An Enterprise AI Delivery Assessment defines that acceptance boundary for a proposed use case, together with the architecture, security controls and responsibilities needed to deliver it.