AI in practice / Colleague-assisted servicing
An AI agent drafts the reply. A colleague decides what happens.
A member asks about their pension transfer. A working agent reads the case and the guidance, then drafts a reply. It cannot change, send or approve anything. Run the requests below and watch where each one stops.
Run the agent
- Member message
- Safety screen
- Agent
- Read-only tools
- Checks
- Colleague review
- Follow-up task
Pick a request above
Each one runs for real. The line above shows how far the request got.
Draft reply to the member
How this run was checked
These rule checks are a starting point. They do not replace a colleague’s judgement.
Every step the backend took
The stored run record
Tasks you approved
Created and stored by the backend. Refresh the page: they are still here.
Try other requests or write your own
Use fictional text only. The task is limited to a transfer-status enquiry; anything else is handed to a colleague.
Why it can be trusted with this
It can only read
The agent has two tools: read this member’s case, and read approved guidance. There is no tool to change a record, send a message or move money, so a manipulated agent has nothing harmful to do.
Facts come from the record
The status and dates in a reply are written by the system from the case record. The model chooses what to read and how to route the request; it does not supply the facts.
A person decides
Every reply is a draft. A colleague approves or rejects it, and only an approval creates a follow-up task. If any check fails, the draft is withheld.
How it is built
Everything runs in one AWS Region (London). The agent is application code in a serverless function; the model and guardrail are Amazon Bedrock services.
How well it works
Small, synthetic test sets run with live model calls from a development machine. They show how this build behaved; they are not a guarantee.
For engineers
Security design
Prompt injection is treated as a risk that remains, not a solved problem. The design limits what a manipulated model could do.
- Guarded inference
- Every model call carries a pinned, numbered Bedrock Guardrail in London. The member’s text, what the tools return and the final draft are screened separately, because the guardrail on a model call does not assess tool results or tool arguments. An intervention, a missing assessment or a screening error withholds the draft.
- Untrusted model output
- Tool names and arguments chosen by the model are validated against a schema. Case ownership comes from the server-side session, never from the model or the browser. An unknown tool ends the run.
- Sessions
- Short-lived demo sessions in a secure, HttpOnly cookie with CSRF protection. They authorise a walkthrough of fictional data; they are not customer authentication.
- Approval
- The backend approves only the exact stored draft: version, hash and text must match. Stale, edited, expired and duplicate approvals are rejected, and concurrent duplicates create one task.
Limits, cost control and operations
- Model choice
- This agent runs on Amazon Nova Pro: the lowest-cost model that passed the evaluation for a routing and tool-calling task. The customer-facing agent on the self-service page runs on Claude Sonnet 4.6. Each is one setting.
- Model gateway
- One model, one Region and one guardrail version, set in code and in the runtime role’s policy. At most three model turns and four tool calls per run, with token, time and context ceilings.
- Spend
- Per-session and daily budgets are checked before every model and guardrail call, including retries. If the counters are unavailable, fresh inference is refused. An operator switch stops inference for the next call, even mid-run.
- Telemetry
- Structured logs record run identifiers, versions, tool steps, outcomes, tokens and latency. Prompts, member text, model output, cookies and tokens are not logged. Alarms cover failures, latency, interventions and exhausted budgets.
Connecting to the backend…
What is real and what is illustrative
Real: the agent loop and its tool calls, the sessions, the stored drafts, the approval check, the tasks, the budgets and the off switch. Each run states whether it used a live model call, a local stand-in or a recorded example. Illustrative: the provider, guidance, cases and follow-up tasks are fictional, and demo sessions stand in for identity.
Public UK pension guidance informed the transfer and advice boundaries: MoneyHelper on pension transfers and on benefits and guarantees.
Want agents like this in your organisation?
I take AI agents from a business problem into production, working with the team that owns it. Tell me what you are trying to do.
Email Dan Message on LinkedIn More use cases