Case study · AI agent supervision · Concept

Watchtower: supervising AI voice agents on live calls

AI voice agents now take real actions during phone calls. They release records, send refunds and order refills. Watchtower lets one person supervise a whole team of these agents and step in before a risky action goes through.

Role

Product design and design engineering

What I built

A supervision dashboard and a call console

Built with

Hand coded HTML and JavaScript

01 · The problem

takeover only

Most agent tools only offer a takeover button

AI voice agents can now act on their own. On a clinic's phone line, that means releasing medical records, sending refunds and requesting refills. A mistake can be expensive: a charge over the limit, health details shared without consent, or a promise the clinic can't keep.

Most products give a person one control: a button to take over the call. That tells you a person can step in. It doesn't tell you when.

What a supervisor needs to know

What is the agent about to do? Is it allowed to do that? Can I stop it in time, without the caller noticing?

Asking a person to confirm every action doesn't work either. An agent takes many actions in a few minutes, and people soon start approving without reading.

02 · The approach

Stop the risky action before it happens

Each agent works within a set of limits, called its scope. Most of the time it acts on its own. When it tries to do something outside those limits, Watchtower pauses that one action and asks a person to decide.

If no one answers

The action does not go through. The agent tells the caller that someone will follow up. Silence never approves anything.

03 · The dashboard

One person, fourteen live calls

The dashboard shows every agent on the line on one screen. It is drawn in the Wolf-Rayet design system, which has one rule: only one thing on a screen may demand attention at a time. Here, that one thing is the decision that needs you now.

Now

one at a time

The decision that needs you, with a countdown. If a second one arrives, it waits underneath.

Live calls

ranked

Every call, sorted by what needs attention soonest. Calls that are going fine are grey, so they don't compete for your eye.

Capacity

wait times

Shows when callers are waiting too long, and whether to add agents or ask for a second supervisor.

Rules

fewer decisions

When you make the same decision a few times, Watchtower offers to make it the agent's standard answer. You can undo it at any time.

Two checks

fewer false alarms

A call only interrupts you when two separate checks agree it needs a person. You can mark any decision as a false alarm.

Catch-up

what changed

When you close a call, a short summary shows what happened on the rest of the line while you were busy.

Live demo · The supervisor's dashboard · Synthetic data

interactive

Try it: make a decision in Now, open one of the four live calls, or add agents. Open in a new tab

04 · The business case

What this changes for a business

A takeover button alone can't make it safe to let agents handle money or medical records. Watchtower is built to make that safe. These are the three things to measure.

01 · More work for the agent

the main gain

Without a way to stop risky actions, agents aren't trusted with payments or records, so people keep doing that work. With Watchtower, agents can take it on within limits a person controls.

Measure: the share of sensitive actions agents are allowed to handle, before and after.

02 · Fewer costly mistakes

loss avoided

Every mistake that is stopped is a cost that never happens: no chargeback, no refund dispute, no privacy complaint. Each one is logged.

Measure: the cost of one incident, times how often it happens, times call volume.

03 · More agents per supervisor

agents per person

If a person has to confirm everything, they can barely keep up with one agent. If they only step in at the limits, they can supervise many. Rules raise that number over time.

Measure: how many agents one person can supervise safely, and how often they are interrupted each hour.

05 · The call console

Deciding on a single call

Open any call and you see the console: the live transcript, what the agent is allowed to do, and the action it wants to take. The supervisor has four choices.

  • Approve the action as it is.
  • Edit it so it fits within the limits.
  • Hold the call. The caller hears a natural pause, not silence.
  • Take over and talk to the caller directly.

If the action is outside the limits, approving it counts as an override. It needs a reason, and it is logged under the supervisor's name.

Example 1 · A records request over the limit · Synthetic data

interactive

Try it: a patient asks for six years of records. The agent may release twelve months. Open in a new tab

06 · Limits and records

Limits in view, and a record of every decision

Limits

always visible

The agent's limits sit next to the transcript, not on a settings page. When the agent reaches one, it changes state right away, so the supervisor can see what the agent is pushing against.

Records

saved automatically

Every decision is saved with a time stamp: what was held, what was approved, what was overridden, and who took over. The record comes from using the tool, so nobody has to write a report after the call.

07 · Consent

When the caller withdraws consent

A supervisor can override the agent's limits. A supervisor can't override the caller's consent.

In this example, a patient asks for a lab result. Then she stops the agent. She is at work, the line might be recorded, and she doesn't want her health details read out.

As soon as she says so, the agent loses permission to share health information. The disclosure is blocked, and there is no approve button. The supervisor can ask for consent again, send the result to her secure patient portal, or take over. Even a person who takes over can't read the result aloud.

Example 2 · Consent withdrawn during the call · Synthetic data

blocked, no approve

Try it: a patient asks for a lab result, then withdraws consent. Open in a new tab

08 · States and failures

Seven states, and what happens when things go wrong

I designed the hard moments first, because that is when a supervisor needs to trust the tool.

01 Watching

normal

The agent is working normally. Nothing needs you.

02 Elevated

near a limit

The agent is getting close to one of its limits.

03 Pending

needs a decision

The agent wants to do something that needs approval. The action waits.

04 Held

paused

You paused the agent. The caller hears a short pause, not silence.

05 Takeover

you're live

You are talking to the caller. The agent is muted.

06 Handback

short brief

The agent takes the call back with a one line summary of what you agreed, so it doesn't contradict you.

07 Resolved

saved

The action went through or was stopped, and the record is saved.

When things go wrong

designed first

No one answers in time.

The action is stopped, not sent. The agent tells the caller that someone will call back.

The action is outside the limits.

It can't go through unless a supervisor overrides it, and the override is logged.

The caller withdraws consent.

Permission to share health or payment details ends right away. No one can override it.

You take over in the middle of a sentence.

The agent finishes its phrase first, so the caller doesn't hear a sudden cut.

The connection to the agent is weak.

Approve and edit turn off, because the command might not arrive. Take over still works by phone.

Two calls need you at once.

Only one takes the Now slot. The other waits right below it, so it can't be missed.

What I decided against

not used

Asking a person to confirm every action

It sounds safe. In practice, an agent makes many decisions in a few minutes. The requests pile up, and people start approving without reading. That is worse than having no check at all.

Watchtower does the opposite. The agent works on its own inside its limits. A person is only asked when it reaches one, with the action explained, a clock running, and a safe result if no one answers.

Explore more

Related work

An incident console with the same visual style, an accessibility tool with the same kind of audit trail, and an interview score that shows its reasoning.

Vantage, finding the cause of an outage
When one fault sets off hundreds of alerts, Vantage traces them back to the one that started it, shows the path it took through the system, and points to the slowest step.
Conformly, checking a site against accessibility law
Scans a website against the European accessibility standards (EN 301 549 and WCAG), opens tested code fixes as pull requests, and records every step. It never says a site is compliant, because only a legal review can say that.
Criterion, interview practice with scores you can question
Every score links to the exact words that earned it. If you disagree with a score, you can challenge it and the AI explains its reasoning again. Live at criterionscore.com.

One person supervising many AI agents, and stepping in before a risky action goes through.

View all case studies

Watchtower is a concept exploration by David Paterni. Interfaces and data are synthetic.