Skip to content
Expert Cloud & AI
Menu

Giving AI agents bounded tasks—and results you can verify

Use structured briefs, measured inputs and checked outputs to make agent research useful, inspectable and safe to act on.

Chris Baran

· 5 min read

On this page
  1. Separate what you measured from what the model found
  2. Write the brief as a bounded research job
  3. Accept a result through a narrow interface
  4. A valid source URL is the beginning of review
  5. Test the failure cases that could change a decision
  6. From repository inventory to useful research

A research agent can produce a convincing answer before anyone has established what it was allowed to assume. That becomes a delivery problem when the output influences a tooling purchase, an architecture decision or a security recommendation.

The useful design starts earlier: define the question, supply measured inputs, bound the research and specify how its result will be accepted. The agent can then do substantial work without becoming the authority for every fact it encounters.

We built a repository-research workflow for an oil and gas company that separates estate analysis, research and reporting. A dispatcher writes the research brief and ingests the returned results. A person or an external agent runs the research task, keeping execution separate from the tool that prepares and accepts the work.

Separate what you measured from what the model found

A repository inventory can establish facts through code: which files were observed, which build definitions were present and which signals an analyser extracted. Research can add a different kind of information, such as whether a product documents support for an integration or how a published limitation affects a use case.

Those inputs deserve different treatment. Preserve measured facts separately from generated research. Record where each research claim came from and what still needs confirmation. An agent should not be able to overwrite a measured result by including a more optimistic value in its response.

100%
Measured repository facts and an approved research brief feed separate paths. Agent output passes through contract checks and human source review before contributing to a decision.

Reference workflow: measured facts remain separate from generated research, and acceptance includes checking the evidence.

Reference workflow: measured facts remain separate from generated research, and acceptance includes checking the evidence.

This distinction is particularly useful in an enterprise assessment. A model can help investigate options while the underlying inventory remains reproducible.

Write the brief as a bounded research job

The workflow divides the work into estate analysis, vendor research and reporting. The brief writer specifies the research question, source expectations and structured output. That makes the assignment portable: a person or an agent can run it, and the returned artifact can be reviewed separately.

A fictional brief might look like this:

A research task with an explicit boundary
task_id: export-format-review
question: Does the documented export preserve record timestamps?
inputs:
  - approved product documentation
  - a synthetic export fixture
return:
  - answer with limitations
  - source URL and title for each supporting reference
  - unanswered questions
rules:
  - do not upload repository contents or credentials
  - do not install tools or change project configuration
  - report missing evidence instead of filling the gap
acceptance:
  - reviewer checks the cited sections
  - fixture results remain separate from the research answer

The surrounding tools must enforce the access boundary. A prompt saying “do not upload” is useful instruction, but restricting the agent’s available files and network destinations is what limits its opportunities to do so. External documents and returned web content are evidence to inspect, not instructions that may expand the task.

Accept a result through a narrow interface

Treat the returned JSON as untrusted input. A small result contract can require the expected task identifier, a bounded summary and references from an admitted source set. Unknown fields should be rejected rather than quietly merged into the authoritative record.

The new companion example makes that admission step explicit:

Accept research without changing measured inputs
import { acceptResearchResult } from "./controls";
 
const measured = Object.freeze({ fixtureChecksPassed: 8 });
const permittedSources = new Set([
  "https://docs.example.org/export-format",
]);
 
const research = acceptResearchResult(
  JSON.parse(agentOutput),
  "export-format-review",
  permittedSources,
);
 
const reviewItem = { measured, research, status: "needs-review" };

Here agentOutput is the response string from the externally run research task. The surrounding runner should also bound response size before parsing. The validator rejects a different task ID, extra fields, missing or duplicated sources, oversized text and source URLs outside the permitted set. It makes no network requests. The permitted set must come from the trusted research workflow, not from a second field invented by the same response.

This illustrative validator checks each returned value before accepting the result. You can download the example and run its tests to see which inputs it accepts and rejects.

A valid source URL is the beginning of review

Admission to the permitted source list says where the response may point. It does not establish that the page supports the summary, that the information is current or that it applies to the proposed environment.

A reviewer still needs to inspect the cited passage, check the date and distinguish a documented capability from an assumption about how it will work. Where a decision matters, use a focused experiment or fixture to test the relevant behaviour.

The same principle applies to retrieval-augmented applications. Scope retrieval to material the requester may access before asking a model to use it. Keep request context separate, and test that a user cannot retrieve another user’s restricted material. Metadata filtering can help select relevant documents, but relevance filtering alone is not an authorisation system. AWS documents retrieval filters as query configuration.

Test the failure cases that could change a decision

For the illustrative contract, useful tests include an output for the wrong task, an unexpected field that tries to replace a measured result, an unapproved source, and a duplicate citation dressed up as additional evidence. A valid response must also pass, so the test suite checks the usable path as well as the rejection path.

Some failures belong above the schema layer. A plausible but unsupported claim can be perfectly valid JSON. A research task can also finish without answering its question. The workflow needs an explicit “insufficient evidence” outcome, with human review before a consequential recommendation is accepted.

That is how the structure creates useful room for autonomy. The agent can work through a substantial research assignment. It returns to a defined interface, and the person making the decision can see what was measured, what was generated and what remains uncertain.

From repository inventory to useful research

The work began with a repository-analysis tool. A later phase added structured research briefs and result ingestion, connecting the inventory to a broader assessment workflow. That gave us a consistent view of the estate to use when investigating options, with sourced findings brought back for review.

The practical gain is a repeatable handoff: the tool prepares a bounded assignment, the agent investigates it, and a reviewer receives structured findings alongside the original measurements. Research can move forward without losing track of which facts came from the repositories and which came from the model.

For your own use case, start with a decision worth supporting and a result you can inspect. Our Enterprise AI Delivery Assessment can define the data boundary, agent task, acceptance checks and ownership before you invest in the wider workflow.

The next article looks at the environment in which that experimentation can happen, including the difference between a configured boundary and one demonstrated by a deployed check.