Algorithmic clarity

Personalization Tuning Studio: make the algorithm legible

See why a title surfaces. Watch whether personalization is healthy. Safely simulate a change.

A ranked row of titles that reshuffles overnight, with no explanation, is a black box, and a black box is impossible to debug, defend, or improve. The Studio is built the other way around: every placement traces back to the exact behavioral signals that earned it, the health of the whole row is observable rather than assumed, and a weighting change can be simulated against a live recommender and watched as the ranking reorders. The one thing the tool will never do is ship that change. The only forward action is to stage it as an experiment for human review. The whole product is a single argument made in interface form: an algorithm people can read, trust, and safely test.

Fully interactive. Move through the whole workspace: Overview → Explain → Health → Simulate → Impact, then drive the signature simulator: change the signal weights and watch the ranking reorder live. Opens in desktop by default; use the toggle for the mobile layout. All data is fictional and original; the sandbox never affects a live recommender.

The problem: a ranked list that moves without explanation

Recommendation systems usually expose a bare ordered list, positions with no traceability and no recourse when something looks wrong.

When a title drops three places overnight, the people responsible for that row have almost nothing to work with. They can't see which behavior moved it, they can't tell whether the model is reacting to a real signal or to noise, and they certainly can't try a fix without pushing code at the live experience. So the row becomes something to fear rather than tune: every change is high-stakes, every explanation is a guess, and "the algorithm did it" becomes an acceptable answer. The first design decision was to treat that opacity as the bug, to make the ranked list into something you can read, diagnose, and rehearse against.

The black box · what most tools expose

  • 1The Quiet Quarter
  • 2Saltwater Kings
  • 3Midnight Atlas

A position, and nothing else. No evidence, no confidence, no way to ask why it moved, and no safe way to try a fix. An ordering you either accept or fight.

The glass box · what the Studio shows

Midnight Atlas · Ranked #3 Explained

In row "Because you watched documentaries" · Confidence: High · 12,400 plays

Completion+ strong
Skip rate− holds back

Every placement traced to the behavioral signals that earned it, and a way to test a change in the simulator.

The thesis: make the algorithm legible and accountable

Three properties, in order. You can explain any placement. You can diagnose the whole row's health. You can safely test a change.

Everything in the Studio descends from one idea: a recommendation people can read, trust, and safely test. A ranking that can't be explained and can't be rehearsed against isn't a tool. It's a verdict the team has to live with. So the interface becomes the mechanism of accountability rather than decoration on a model: each placement is shown with the signals behind it, the health of personalization is observable rather than assumed, and any change is run in a sandbox first. Trust here isn't a reassuring tone of voice. It's a structural property you can poke at.

Explain a placement

Pick a title and a segment and see why it's ranked where it is: completion, hover, skip, recency, and taste match, each shown as a contribution you can read. Nothing is ranked against something you can't see.

Diagnose health

Watch rank-over-time and volatility for a title or segment, with a plain verdict: Healthy, Stabilizing, or Worth a human look, and a timeline of what actually changed.

Safely test a change

Move the signal weights and watch the ranking reorder live, with predicted impact, then stage it as an experiment for review. The tool never ships; a person decides whether it ever runs.

The core decision: a ranking you can rehearse, not ship

The simulator isn't a settings panel. It's the interaction that turns a feared, high-stakes change into a rehearsed, reviewable proposal.

Each title carries behavioral signals scored 0 to 1. The ranking is a weighted sum: completion, hover, skip penalty, recency, and a diversity bonus, and the sliders set those weights. Move one and the Studio recomputes every score, re-sorts the row, animates the reorder, and marks each move with a ▲ / ▼ / ─ delta against the frozen production baseline. A predicted-impact panel updates live: engagement, a diversity index, and a member-experience risk read. But there is deliberately no "Ship" button anywhere. The only forward action out of a simulation is Stage as A/B experiment, which routes the weights and predicted impact to experimentation review. A person decides whether it ever reaches members.

Production baseline · frozen

  • 1The Quiet Quarter0.81
  • 2Saltwater Kings0.77
  • 3Midnight Atlas0.74
  • 4Origin Unknown0.69

Read-only. This is what members see today; it never moves while you experiment.

Simulation · default weights → reordered

Simulated
  • 1Midnight Atlas▲2◆0.83
  • 2The Quiet Quarter▼1◆0.80
  • 3Saltwater Kings▼1◆0.76
  • 4Origin Unknown◆0.70

Lifting completion weight moves Midnight Atlas ▲2 to #1. Predicted impact updates live: engagement ◆+2.1%, diversity index ◆0.71, member risk ◆Low/Med, and every simulated value carries the ◆ mark. Reset to production baseline is the in-session undo.

Stage as A/B experiment Routed to review Reset baseline

This is the signature interaction. Run it yourself in the prototype above. Simulate resets to the production baseline each visit, so the flow is always replayable.

Honesty about uncertainty is the trust-builder

Confidence levels, provisional signals, and a "worth a human look" verdict are shown on purpose. The Studio states what it isn't sure of rather than hiding it.

A recommender is never uniformly confident, and pretending otherwise is how trust dies. When a signal is thin, an artwork CTR computed on fewer than 500 impressions, the Studio marks it provisional instead of folding it silently into the score. When a title's rank has been swinging, Health doesn't paper over it; it routes a plain-language verdict, Worth a human look, with the confidence stated as Low. Confidence is a first-class label, not a hidden parameter. Admitting what's uncertain is exactly what makes the confident placements believable.

State = shape first, colour second

Explained: traced to signals Provisional: thin evidence Awaiting: not yet measured Simulated: sandbox, never live Worth a human look: routed to review

Provisional signal · flagged, not folded in

Artwork CTR Provisional

"Computed on fewer than 500 impressions, shown so you can weigh it, not hidden so it quietly tilts the rank." A signal the Studio refuses to treat as settled.

The hardest decision: a wall between production and the sandbox

Production and Simulation sit side by side, always. Every simulated value carries the ◆ mark. The boundary is spatial, not a label you have to trust.

The moment a tuning tool can silently change what members see, it stops being safe to explore in, and an unsafe tool gets used timidly or not at all. The load-bearing decision was to make the wall physical: a read-only PRODUCTION column on one side, a live SIMULATION column on the other, never collapsed into a single editable view. Nothing crosses that wall without a human. And the Studio is honest about cost on both sides of it. When a weighting change lifts engagement but drops the diversity index, the panel says so plainly rather than celebrating the win and hiding the trade. "Watch this" is a feature, not a confession.

Production · read-only

What members actually see right now.

Frozen while you experiment. The Studio cannot write to it. The only path here is through experimentation review, decided by a person.

Simulation · sandbox

Where every change lives until a human decides.

Recomputes live, marks every value with ◆, and never touches the live recommender. Reset to production baseline at any time.

Diversity index ◆0.71 −0.04 vs baseline · watch this

The same change that lifts engagement narrows the row slightly. The Studio surfaces the trade-off next to the win, so the person reviewing the experiment sees the cost, not just the headline.

"Safely test," made literal: a simulation becomes a proposal a person reviews

Staging an experiment doesn't change members' experience. It produces a clean before/after record and routes it to experimentation review.

Tuning a ranking is good. Tuning it into an artifact someone can actually decide on is better. When you stage a simulation, the Studio assembles a proposal: the from/to on every weight, the resulting #1 title, the predicted engagement and diversity, and the honest risk read, and stamps it Routed to experimentation review. It does not change what members see; a person decides whether it ever runs. That round trip, from a moved slider to a reviewable proposal, is the most concrete answer to "what does safely test mean?"

Proposal · lift completion weight on the documentaries row

Routed to review
Completion weight30 ◆40
Skip penalty weight15 ◆20
#1 titleThe Quiet Quarter ◆Midnight Atlas
Diversity index0.75 ◆0.71 watch this

This is a static specimen. Stage a simulation in the prototype above to generate the live proposal and confirm sheet.

v1 → v7: sharpening one thesis, not stacking features

Each version defends the same idea: make the algorithm legible and safe to tune, against the easier, less honest version of every feature.

The arc isn't a changelog. Early versions found the product identity and put the simulator at the centre. The middle versions made the math real and built the explanation layer that keeps a placement honest. The late work is where the thesis got defended: a hard production-versus-sandbox wall, honest trade-off reporting, a Material Design 3 system that holds up in light and dark, and accessibility done as real contrast math rather than a checkbox.

  • v1

    Component dump.

    A grayscale dashboard of charts. It looked like an analytics tool, not a product, and there was no thesis yet, just numbers.

  • v2

    Found the thesis.

    "Make the algorithm legible." The five-screen spine (Overview, Explain, Health, Simulate, Impact) and the simulator taking centre stage as the signature interaction.

  • v3

    Real weighted math.

    Signals scored 0 to 1, a genuine weighted sum, a live recompute and FLIP reorder with ▲/▼ deltas against a frozen baseline, not a canned animation.

  • v4

    The explanation layer.

    Per-placement contribution bars (positive filled right, negative outlined left, never colour alone), a copy-to-clipboard plain-language explanation, and "see the evidence" expanders that show the numbers behind each signal.

  • v5

    Honest uncertainty + the no-ship law.

    Provisional and Worth-a-human-look states; confidence as a first-class label; and the load-bearing rule: no "Ship" button anywhere, only "Stage as A/B experiment" routed to human review behind a confirm sheet.

  • v6

    The production↔simulation wall.

    Made the sandbox boundary spatial: read-only Production beside live Simulation, every simulated value marked ◆, "reset to production baseline" as in-session undo, and trade-offs (the diversity dip) surfaced next to the wins instead of buried.

  • v7

    Material Design 3 + accessibility, done properly.

    A full M3 system: navigation rail that becomes a bottom bar, tonal surface elevation, state layers, the signature sliders, and a type/shape scale, with a warm-cool azure palette over the default purple. Dark mode's real failures were accent-as-text colours under the 4.5:1 AA line, fixed with dedicated text-only variables that brighten in dark; the FLIP reorder and reveals honor reduced-motion.

Explore more work

More explorations from the AI Product Design Lab, each a different facet of making AI products people can direct, verify, supervise, and trust.

steer exploration cover
Steer, intent before generation

Turn an under-specified prompt into a negotiated brief: the model surfaces what it inferred and flags ambiguity before it commits.

View exploration
ground exploration cover
Ground, verify what AI claims

Every claim traceable to a source with confidence and freshness; unsupported claims flagged; source conflicts shown, not smoothed over.

View exploration
recall exploration cover
Recall, legible AI memory

A memory layer you can see, attribute, edit, scope, and revoke, personalization as a negotiated, inspectable thing, not a black box.

View exploration