Forward Deployed Product Manager

9. Delivery, Triage & Recovery

Operate multiple deployments and recover failing AI systems.


9.1Multi-account triage (link to this section)

Accounts A and B conflict. Decide who gets today and defend the decision.

Multi-account triage

Two accounts need you today. Only one can have you.

The Core Idea

Triage requires a defensible framework — contract value, churn risk, issue severity, strategic importance — rather than defaulting to whichever customer emails loudest.

The Defensible Framework

Weighing contract value, risk of churn, severity, and strategic importance consistently means the decision can be explained the same way every time, not reconstructed after the fact to justify whichever choice already felt easiest.

Explaining the Decision to the Deprioritized Account

A clear "not today, and here's why, and here's when" — grounded in the same criteria used every time — preserves the relationship far better than silence or an unexplained delay.

Where This Shows Up in the Field

This is Section 3's ICP-as-capacity-allocation logic under live time pressure — the rubric built there for choosing accounts is the same rubric that should decide which account gets today.

9.2Blocker diagnosis (link to this section)

Determine whether the blocker is product, execution, integration, data, security or change management.

Blocker diagnosis

"We're stuck" isn't a diagnosis. It's a symptom.

The Core Idea

A stalled deployment could be blocked by the product, execution capacity, integration complexity, data issues, security review, or organizational resistance. Each has a completely different fix.

Why Category Matters More Than Speed

Ask specifically: is this blocked because the product can't do it, because it hasn't been built yet, because a system won't connect, because data isn't there, because security hasn't approved it, or because people won't use it. Each answer points to a different owner and a different fix.

Escalating to the Wrong Owner Wastes Time Twice

Misdiagnosing category means escalating to someone who can't actually resolve it, then re-diagnosing and escalating again — slower overall than taking the extra minute to categorize correctly first.

Where This Shows Up in the Field

This connects directly to Section 8's organisational resistance lesson — "we're stuck" that's actually a resistance problem needs a completely different response than "we're stuck" that's a genuine integration blocker.

9.3AI failure diagnosis (link to this section)

"The answers are wrong" — determine whether the cause is retrieval, chunking, prompt, context, tool or model.

AI failure diagnosis

"The answers are wrong." Now what?

The Core Idea

The same diagnostic chain from Section 5's RAG architecture lesson — retrieval, chunking, prompting, context, tools, model — applied under live production pressure, where the instinct to guess fast is strongest.

The Instinct to Resist

Under customer pressure, guessing at a fast fix feels productive. It usually isn't — a wrong guess wastes time and may not address the actual cause, while the same methodical chain that works offline still finds the real issue fastest, even under pressure.

1Trace inspection
2Retrieval quality
3Prompting
4Context
5Tools
Resist guessing — work the chain live

(Decision tree)

Working the Chain Live

Trace inspection first, then retrieval quality, then prompting — the order doesn't change just because the customer wants an answer immediately.

Where This Shows Up in the Field

This is Section 5's RAG diagnosis tree and Section 6's incident response discipline, applied simultaneously, under the added pressure of a customer watching in real time.

9.4Post-launch iteration & recovery (link to this section)

Tune against real usage and rescue a failing deployment. Ship: retro + fix plan.

Post-launch iteration & recovery

Launch is the start of the real data, not the finish line.

The Core Idea

Real usage patterns after launch almost always differ from what discovery and testing anticipated. Tuning against actual production data is as core to the role as the initial build.

Why Post-Launch Data Surprises Even Good Discovery

Discovery and offline evaluation (Section 6) test against a fixed, known set of cases. Real usage generates cases no one anticipated — this is expected, not a sign the earlier work was wrong.

Recovery Requires the Same Discipline as Diagnosis

A genuinely failing deployment needs the same methodical diagnostic approach as any other AI failure — panic and quick, unverified fixes tend to make a bad situation harder to reason about, not better.

Where This Shows Up in the Field

A retrospective without a committed fix plan doesn't actually recover anything — diagnosis alone is only half the job.

Ship: Retro + Fix Plan

Pair an honest retrospective (what went wrong and why) with a concrete fix plan — specific changes, owners, and timeline — not just a diagnosis without a committed path to resolution.

Practise this chapter in the workspace

Reading is the map. Every section above also runs as a hands-on workspace session with tools, exercises and a recap quiz.

Start Learning for Free