AI in practice / Colleague-assisted servicing

An AI agent drafts the reply. A colleague decides what happens.

A member asks about their pension transfer. A working agent reads the case and the guidance, then drafts a reply. It cannot change, send or approve anything. Run the requests below and watch where each one stops.

AnyCompany Pensions is fictional and all data is synthetic. This is a demonstration, not a pension service.

Run the agent

  1. Member message
  2. Safety screen
  3. Agent
  4. Read-only tools
  5. Checks
  6. Colleague review
  7. Follow-up task

Pick a request above

Each one runs for real. The line above shows how far the request got.

Try other requests or write your own

Use fictional text only. The task is limited to a transfer-status enquiry; anything else is handed to a colleague.

Why it can be trusted with this

It can only read

The agent has two tools: read this member’s case, and read approved guidance. There is no tool to change a record, send a message or move money, so a manipulated agent has nothing harmful to do.

Facts come from the record

The status and dates in a reply are written by the system from the case record. The model chooses what to read and how to route the request; it does not supply the facts.

A person decides

Every reply is a draft. A colleague approves or rejects it, and only an approval creates a follow-up task. If any check fails, the draft is withheld.

How it is built

Architecture of the demonstration A browser calls an API gateway in London, which calls the agent runtime. The agent runtime uses a guardrail, a language model and two read-only tools. A colleague’s approval is checked against the stored draft before a task is created. BrowserAPI gatewayAgent runtime GuardrailLanguage modelRead-only tools ColleagueApproval checkTask store holds no credentialsLondon, HTTPS only bounded loop: 3 model turns,4 tool calls, spend limits,operator off switch screens the message, what thetools return and the final draft chooses tools and routes therequest; does not supply facts this member’s case record,approved guidance approves or rejectsexact stored draft onlyone task per approval stored draft

Everything runs in one AWS Region (London). The agent is application code in a serverless function; the model and guardrail are Amazon Bedrock services.

How well it works

Small, synthetic test sets run with live model calls from a development machine. They show how this build behaved; they are not a guarantee.

For engineers

Security design

Prompt injection is treated as a risk that remains, not a solved problem. The design limits what a manipulated model could do.

Guarded inference
Every model call carries a pinned, numbered Bedrock Guardrail in London. The member’s text, what the tools return and the final draft are screened separately, because the guardrail on a model call does not assess tool results or tool arguments. An intervention, a missing assessment or a screening error withholds the draft.
Untrusted model output
Tool names and arguments chosen by the model are validated against a schema. Case ownership comes from the server-side session, never from the model or the browser. An unknown tool ends the run.
Sessions
Short-lived demo sessions in a secure, HttpOnly cookie with CSRF protection. They authorise a walkthrough of fictional data; they are not customer authentication.
Approval
The backend approves only the exact stored draft: version, hash and text must match. Stale, edited, expired and duplicate approvals are rejected, and concurrent duplicates create one task.
Limits, cost control and operations
Model choice
This agent runs on Amazon Nova Pro: the lowest-cost model that passed the evaluation for a routing and tool-calling task. The customer-facing agent on the self-service page runs on Claude Sonnet 4.6. Each is one setting.
Model gateway
One model, one Region and one guardrail version, set in code and in the runtime role’s policy. At most three model turns and four tool calls per run, with token, time and context ceilings.
Spend
Per-session and daily budgets are checked before every model and guardrail call, including retries. If the counters are unavailable, fresh inference is refused. An operator switch stops inference for the next call, even mid-run.
Telemetry
Structured logs record run identifiers, versions, tool steps, outcomes, tokens and latency. Prompts, member text, model output, cookies and tokens are not logged. Alarms cover failures, latency, interventions and exhausted budgets.

Connecting to the backend…

What is real and what is illustrative

Real: the agent loop and its tool calls, the sessions, the stored drafts, the approval check, the tasks, the budgets and the off switch. Each run states whether it used a live model call, a local stand-in or a recorded example. Illustrative: the provider, guidance, cases and follow-up tasks are fictional, and demo sessions stand in for identity.

Public UK pension guidance informed the transfer and advice boundaries: MoneyHelper on pension transfers and on benefits and guarantees.

Want agents like this in your organisation?

I take AI agents from a business problem into production, working with the team that owns it. Tell me what you are trying to do.

Email Dan Message on LinkedIn More use cases