Ship — discovery, scoring, RAID, change, value
The AI delivery playbook (your actual job)
Ship like a delivery lead
Why this matters for a delivery manager
This is the day the rest of the month exists to serve. You already know delivery. Today you write the variant that hiring managers for AI Delivery, FDE, and Transformation will ask you to perform on a whiteboard.
Discovery that ends in 'they want a chatbot' is a failed meeting. Discovery that ends in a job-to-be-done, a corpus note, an eval definition, a score, and a no-AI alternative is a meeting you can fund or kill. Scoring is how you stop HIPPO from commissioning twelve POCs. RAID is the same template with new rows. Change is trust plus a roster, not a launch email.
You will be tempted to skip the no-AI alternative because it feels disloyal to the brief. Write it anyway. If a form or a search improvement would do the job, the honest delivery lead says so. That sentence is worth more in an interview than a clever agent diagram.
You will be able to
- Run a 90-minute discovery that ends in a scored use case, not a chatbot request
- Apply a scoring model: value, feasibility (data/evals/ACL), risk, change
- Write AI-specific RAID and a change plan that includes HITL roster and trust
- Sequence shadow → assist → confirm, with kill criteria
2-hour clock
120:00
Now: Read the playbook · 50m
The 2-hour session
Concepts, in full
This block is a slow read — about an hour with the diagrams. After each concept, write one sentence in notes (what you already do vs what is new) and tick annotated. Do not skim the last concept.
01
Discovery is jobs, not models
Ninety minutes. In the room: a user who does the work, a knowledge or system owner, you, optionally an engineer. Out of the room: a vendor. Agenda with times, not a free conversation that becomes a product pitch. Watch or walk the current task (15). Pain and frequency (10). Data where it lives and class (15). What 'good' means in their words (10). Thin-slice options including non-AI (15). Risks they fear (10). Score and next step (15).
Outputs you leave with: a job-to-be-done in one sentence, a corpus note (where, class, owner), a first eval definition in their words, a four-dimension score, and a no-AI alternative — process change, search, a form. If you cannot name a no-AI alternative, you did not understand the job. You understood a request for a chatbot.
Vendors do not belong in discovery. They belong later, on a scorecard, on your set. A vendor in the room will steer toward their object. You will leave with a feature list instead of a job. You already know not to invite a systems integrator to the problem-framing workshop if you want an honest problem. Same rule.
Watch the work if you can. A PM searching Slack, Confluence, and a slide deck for twenty minutes to find a decision is a different problem from 'we need a knowledge assistant.' The twenty minutes is the job. The three systems are the corpus. The definition of good is 'decision with a citation in under two minutes.' You cannot get that from a survey that asks whether people would use AI.
What 'good' means must be said in their words and then turned into an eval row. 'I would trust it if it showed the page' is a citation requirement. 'I would use it if it did not make up owners' is a groundedness and refuse requirement. Write the sentence they said, then the metric. This is how day 16 gets its first ten rows without you inventing them in a cafe.
End with a next step that is a score and a decision, not 'we will explore tools.' Fund a thin slice. Kill. Or send back for data-ready work with an owner and a date. Three exits. Meetings that end in 'this was great, let's keep talking' are how twelve POCs start. You are the person who puts a number on the board before people stand up.
Vendors do not belong in discovery. They belong later, on a scorecard, on your set. A vendor in the room will steer toward their object and you will leave with a feature list instead of a job. You already know not to invite a systems integrator to the problem-framing workshop if you want an honest problem. Same rule. If a sponsor brings a vendor anyway, park them in a later meeting and keep this ninety minutes on the work as it is done today.
Watch the work if you can. A PM searching Slack, Confluence, and a slide deck for twenty minutes to find a decision is a different problem from 'we need a knowledge assistant.' The twenty minutes is the job. The three systems are the corpus. The definition of good is a decision with a citation in under two minutes. You cannot get that from a survey that asks whether people would use AI. Time it. Write it. That sentence becomes an eval row the same afternoon.
Diagram
90-minute discovery
Walk the job (15)
Watch or reconstruct the current task. Tools they actually touch. Time it.
Pain and frequency (10)
How often, how long, who cares when it fails. Value is here.
Data and class (15)
Where it lives, owner, class, ACL, junk and secrets. Feasibility starts here.
Definition of good (10)
Their words, then an eval row. Include what must be refused.
Slices and no-AI (15)
Thin AI slice, and the form/search/process you would ship if models did not exist.
Risks they fear (10)
Audience, leak, wrong advice, job loss. Change and risk scores live here.
Score and next (15)
Four dimensions, a number, fund / kill / send back for data-ready. Written before people leave.
Times are part of the design. Without them, the last 40 minutes become a chatbot wishlist and you leave with no score.
02
Score like a portfolio manager
Four dimensions. Value: frequency times time or money times willingness to change. Feasibility: data-ready, ACL path, a possible thin slice in four weeks, eval-able. Risk: blast radius, regulation, injection and write-tools. Change: a named owner and users who will do HITL. Score 1 to 5 with a one-line why. Publish the rubric. Let people argue with the rubric, not with you.
Kill anything that is high value but not eval-able. You will never know if you won. 'Executive assistant that just gets it' is not eval-able. 'Find the last recorded decision on a workstream in under two minutes with a citation' is eval-able. The second can be high value. The first is a wish that will consume a year. Feasibility includes eval-able. Do not hide that inside 'technical complexity.'
Feasibility is mostly data. If the corpus is not classed, owned, or ACL-mappable, the score is low no matter how pretty the demo. A four-week slice you cannot data-ready is not a four-week slice. It is a four-week wait plus a slice. Score the wait. Sponsors hate this dimension because it feels like obstruction. It is the dimension that predicts whether you go live.
Change is the dimension delivery leads forget to score, then relive in adoption failure. Is there a named owner who will spend hours. Are there users who will sit in shadow mode. Is HITL staffable. If the answer is 'the organization should want this,' the score is 1. Desire at the top with no roster in the middle is how copilots launch to silence.
HIPPO — highest paid person's opinion — is how you get twelve POCs. A published score does not stop a HIPPO from asking. It forces them to override in the open: 'I am funding the 2 despite the rubric because of X.' That sentence is governance. A quiet 'just do the CEO one too' is how the portfolio dies. You will still lose some of these. You will lose fewer, and you will have a paper trail.
Use the scores to sequence, not only to kill. High value, high feasibility, low risk, adequate change: fund now. High value, low feasibility: data-ready workstream, not a model workstream. High risk: shadow only, or kill. Low value: do not fund even if easy. Easy and pointless is how you spend a quarter proving you can ship trivia.
Feasibility is mostly data. If the corpus is not classed, owned, or ACL-mappable, the score is low no matter how pretty the demo. A four-week slice you cannot data-ready is a four-week wait plus a slice. Score the wait. Sponsors hate this dimension because it feels like obstruction. It is the dimension that predicts whether you go live. Change is the other one delivery leads forget: named owner, HITL hours, users who will shadow. 'The organization should want this' is a 1.
HIPPO is how you get twelve POCs. A published score does not stop a HIPPO from asking. It forces an override in the open: 'I am funding the 2 despite the rubric because of X.' That sentence is governance. A quiet 'just do the CEO one too' is how the portfolio dies. You will still lose some of these. You will lose fewer, and you will have a paper trail. Keep the rubric public so the argument is about the rubric, not about you.
| Dimension | 1 | 3 | 5 | Kill if |
|---|---|---|---|---|
| Value | Rare, small time, nobody will change | Weekly pain, real minutes, some willingness | Daily expensive pain, users asking for help now | Value unknown and nobody will help define it |
| Feasibility | No owner, mixed class, cannot ACL, not eval-able | Corpus exists, class-able in two weeks, eval rows possible | Data-ready in sight, thin slice in four weeks, eval-able now | Cannot eval, or cannot data-ready / ACL in a thin slice |
| Risk | Customer / money / safety / write-tools unbounded | Internal write with confirm; some PII | Internal read, citations, no writes | Unbounded write to customers without a path to never-auto |
| Change | No owner, no HITL hours, users hostile | Owner named, hours unclear, a few champions | Owner + roster + users who will shadow | No human will review or champion; launch would be theatre |
Scoring rubric. Publish it. Argument belongs on the rubric, not on the person who applied it.
Diagram
Value × feasibility (risk and change sit on the rows you fund)
| Low feasibility | High feasibility | |
|---|---|---|
| High value | Data-ready workstream, not a model. Re-score after the checklist. | Fund a thin slice. Sequence shadow → assist. |
| Low value | Kill. Do not 'keep it warm.' | Kill or park. Easy trivia consumes the roster you need elsewhere. |
Fund the top-right. Send top-left to data-ready. Do not fund bottom-right trivia. Kill what you cannot eval.
03
RAID, rewritten for models
You already write RAID. The novelty is the rows, not the template. Risks: hallucination on high-cost acts, ACL leak, stale corpus, cost runaway, vendor deprecation, prompt injection, HITL understaffed, adoption failure, secrets in logs, floating model pin. Issues: eval below bar, data-ready red, DPA unsigned. Assumptions: 'users will cite-check,' 'the corpus is current,' 'twelve users represent the firm.' Dependencies: identity provider, knowledge-owner hours, DPA, enterprise endpoint access, labeling time.
Write them as you would any other RAID: owner, likelihood, impact, mitigation, trigger. 'Prompt injection' without a mitigation is a keyword. The mitigation is the stack from day 15 and the HITL dial from day 17. 'Cost runaway' without a trigger is a worry. The trigger is the daily cap from day 18. Tie each row to a control you have already designed. Orphan rows are theatre.
Stale corpus is the risk sponsors underestimate. A wiki copilot that cites a 2022 SOP against a 2025 process is a confident wrong system. Mitigation: corpus owner, recrawl cadence, sample of twenty, a 'last touched' filter if you can. Issue when the owner is on leave and the recrawl is red. You already manage stale documentation on programs. Here the model will speak the stale page with more confidence than the page ever had.
HITL understaffed belongs in RAID, not only in the change plan. If the roster is two people and one leaves, you are in bound-auto you did not design, or you are in a queue that will overflow. Mitigation: backup names, a freeze-to-assist if coverage drops, a named trigger. This is the same as a single point of failure on a cutover. You already flag those.
Assumptions that users will cite-check should be treated as false until proven. Shadow and assist exist partly because they will not. If your mitigation for hallucination is 'the PM will notice,' you have an assumption, not a control. Put it in assumptions, then put a real control in risks. Mixing those columns is how RAID goes green while the design is hollow.
Review RAID on the same cadence as eval, not only at steering. A weekly fifteen minutes: which row moved, which trigger fired, which mitigation is late. AI RAID that is only a charter appendix is a dead document. You know this. Do not let the novelty of the rows make you forget the operating rhythm.
Stale corpus is the risk sponsors underestimate. A copilot that cites a 2022 SOP against a 2025 process is a confident wrong system. Mitigation: corpus owner, recrawl cadence, sample of twenty, a last-touched filter if you can. You already manage stale documentation. Here the model will speak the stale page with more confidence than the page ever had. Put the recrawl on the RAID trigger list so it is an operating item, not a folklore complaint.
Assumptions that users will cite-check should be treated as false until proven. If your mitigation for hallucination is 'the PM will notice,' you have an assumption, not a control. Put it in assumptions, then put a real control in risks: HITL, citations in the UI, groundedness bar. Mixing those columns is how RAID goes green while the design is hollow. HITL understaffed belongs here too, with a freeze-to-assist if coverage drops. Single points of failure on a cutover get flagged. Flag this.
| Type | Row | Mitigation you can actually run | Trigger |
|---|---|---|---|
| Risk | Hallucination on a high-cost act | Confirm or never-auto; groundedness bar; citations required | Eval drop or a SEV on a send |
| Risk | ACL leak through retrieve | Split indexes; identity on retrieve; ACL test rows in golden set | A user sees a page they cannot open in source |
| Risk | Cost runaway | Routing, cache, daily cap with a page | Daily $ alarm |
| Risk | HITL theatre / understaffed | Named roster, hours, freeze-to-assist if coverage drops | Queue overflow or backup on leave |
| Risk | Vendor deprecation / floating pin | Pin, deprecation drill, successor eval | Vendor email or unexplained quality change |
| Assumption | Users will cite-check | Treat as false; HITL and citations in the UI; measure click-through | Thumbs-up on uncited answers |
| Dependency | Corpus-owner hours for data-ready | Named owner and backup; gate the start date | Sample of 20 not done |
| Issue | Eval below bar | Freeze prompt changes; waiver path or rollback | Weekly sample misses the number |
Starter AI RAID rows. Copy the ones that apply; do not paste all of them unread.
04
Change is trust plus a roster
People do not adopt a copilot that was wrong twice in week one with no citation. Trust is earned in order: shadow mode earns traces and a quiet comparison to the old path. Assist mode earns a habit because the human still sends. Champions, office hours, and a visible kill switch earn permission to keep going when it misses. Benefits tracking — time, quality, avoided rework — earns the second year of funding. Skip the order and you will ask for adoption with nothing in the bank.
Who loses a task, and who is afraid they lose a job, belongs in the change plan. If you skip it, you will meet it as resistance and call it 'culture.' A status-email draft that saves twenty minutes is a gift. A copilot that looks like it might replace the PMO's reporting role is a threat. Tell the truth about the job design. v1 is assist. v1 is not a headcount plan. If someone is planning headcount off v1, that is a different conversation and it should not hide inside your charter.
Champions are named users who will do the ugly first weeks and tell you when it is wrong. Three is enough. They need time, not a title. Put the time in their manager's view so it is real. Office hours — thirty minutes twice a week for the first month — catch the cases that would otherwise become Slack mockery. You already run this pattern on system rollouts. Copy it. Add 'bring the bad answer' as the agenda.
The kill switch is a change artifact, not only an ops artifact. Users who know you can turn it off in minutes are more willing to try it. Users who think they are stuck with a bad copilot until next quarter will not try it, or will try it once and campaign against it. Show them the switch. That is part of the launch, not a secret runbook.
Communications should say what it will refuse, not only what it will do. 'It will not answer HR salary, it will not email customers, it will show sources or it will say it does not know.' Those sentences prevent the first wrong expectation, which is the first lost champion. Over-promising 'it knows our company' is how you get the magical mental model from week 1 all over again.
Benefits tracking starts in week one of assist, not at the annual review. Time on a sampled task, rework on status emails, tickets opened that did not bounce. Small n is fine. Direction is what funding needs. A delivery lead who cannot say 'we timed eight tasks and saved four minutes on six of them, with two misses we put in the golden set' will lose the second year to a shinier vendor.
Who loses a task, and who is afraid they lose a job, belongs in the change plan. A status-email draft that saves twenty minutes is a gift. A copilot that looks like it might replace the PMO's reporting role is a threat. Tell the truth about job design. v1 is assist. v1 is not a headcount plan. If someone is planning headcount off v1, that is a different conversation and it should not hide inside your charter. Hidden headcount plans are how champions become opponents.
Champions are named users who will do the ugly first weeks and bring you the bad answers. Three is enough. Put their time in their manager's view. Office hours twice a week for the first month catch the cases that would otherwise become Slack mockery. Communications should say what it will refuse, not only what it will do. Over-promising 'it knows our company' recreates the magical mental model from week 1. The kill switch is a change artifact: users who know you can turn it off will try it.
05
Sequence and kill criteria, published before you start
Default sequence: shadow for two weeks, assist for four, then maybe confirm on one write. Kill if eval or adoption misses the bar. Put that on the charter tomorrow. A sequence without kill criteria is a roadmap that cannot fail, which means it cannot be managed. You already put kill criteria on troubled programs. Put them on healthy ones too, while people are calm.
Shadow exit: recall and groundedness at the bar on the golden set, plus traces from real work, plus a champion who will say out loud whether it would have helped. If you cannot exit shadow, you do not enter assist. Entering assist because the calendar moved is how week one of assist becomes two public misses and a trust hole you will not fill this quarter.
Assist exit to confirm: the same quality bar, plus reject reasons that are not dominated by 'wrong source,' plus a roster that actually reviewed, plus a written envelope for the write. If reject reasons are mostly retrieve fails, you do not have a confirm problem. You have a retrieve problem, and adding a click will not fix it.
Kill criteria examples, written in numbers. Eval: groundedness under 0.8 on two consecutive weekly samples, or must-refuse fail. Adoption: fewer than eight of twelve named users active after three weeks of assist, without a reason you accept. Risk: an ACL leak, or a SEV on a send. Cost: daily cap hit three days in a week. Any one of these pauses. Two of these kill unless a named exec waives with an expiry. Copy the waiver language from day 16.
Publish the kill before go-live so it is not personal later. When you kill, you are executing a plan, not attacking a sponsor's idea. When you pause, you are executing a plan. The difference between a delivery lead and a mascot is whether those sentences were written when everyone still liked each other.
Killing a slice is not killing AI at the company. It is how you keep the next slice fundable. A zombie copilot that nobody uses and nobody will turn off poisons the portfolio. You have seen zombie tools. Do not add one with a model in it. The death date from day 19 and the kill criteria from today are the same instinct applied to POC and to production.
Shadow exit is a gate: recall and groundedness at the bar, traces from real work, a champion who will say out loud whether it would have helped. If you cannot exit shadow, you do not enter assist. Entering assist because the calendar moved is how week one becomes two public misses and a trust hole you will not fill this quarter. Assist exit to confirm needs reject reasons not dominated by wrong source, a roster that actually reviewed, and a written envelope for the write.
Write kill criteria in numbers while people still like each other. Groundedness under 0.8 two weekly samples in a row, or must-refuse fail: pause. Fewer than eight of twelve named users active after three weeks of assist, without a reason you accept: pause. ACL leak or SEV on a send: kill unless waived. Daily cap hit three days in a week: pause. Two pauses without a fix: kill unless a named exec waives with expiry. Copy the waiver shape from day 16. Publish it before go-live so using it is not personal.
06
HIPPO, portfolio, and saying no without theatre
HIPPO is not always wrong. It is always unaccountable unless you make it accountable. When the highest-paid person wants the chatbot of everything, you score it in the open, you show the feasibility 1 and the risk 5, and you offer the thin slice that is adjacent and fundable. If they still override, they override in writing. You run the slice they funded, with the bar, or you refuse if it violates a control you cannot waive (restricted data with no path). Refusing a control violation is the job. Sulking about the override is not.
A portfolio of more than three active AI slices with the same HITL roster is a fantasy. Capacity is math, again. Score, sequence, park. Parking is a decision. 'Later, after Atlas is in assist and the roster has slack' is a date. 'On the backlog' is a grave. You already run intake this way. Do not let AI ideas skip intake because they are exciting.
Saying no without theatre means a short written no: the score, the missing dimension, the door back in ('data-ready checklist complete and we re-score'). A long slide about strategy is theatre. People need the door. Without a door, they will go around you to a vendor POC, and you will meet them on day 19 with a leaked production. Give them a door that is real.
Keep a public board: funded, data-ready waiting, parked, killed. Four columns. Update it when you score. Sponsors can see their idea. You can see the volume. Steering can see that kill is a column that sometimes has cards in it. A board that only shows funded work is how HIPPO assumes everything lives.
Your credibility is the rubric plus one shipped slice. A delivery lead who scores everything to death and ships nothing is a bottleneck. A delivery lead who ships without scoring is a mascot. This week is the operating system. Tomorrow's charter is the first slice. Do both. The playbook that never produces a charter is a blog.
When a hiring manager says 'walk me through how you would intake AI use cases,' this day is the answer: 90-minute discovery, four-dimension score, RAID rows, shadow-assist-confirm, kill criteria, HIPPO in the open. Draw it. Do not talk about frameworks with names. Talk about the meeting, the table, and the death of a bad idea. That is the job.
A portfolio of more than three active AI slices on the same HITL roster is a fantasy. Capacity is math. Score, sequence, park. Parking is a decision with a date: later, after Atlas is in assist and the roster has slack. 'On the backlog' is a grave. Keep a public board: funded, data-ready waiting, parked, killed. A board that only shows funded work is how HIPPO assumes everything lives. Saying no without theatre is a short written no plus a door back in: finish the data-ready checklist and we re-score.
Your credibility is the rubric plus one shipped slice. A delivery lead who scores everything to death and ships nothing is a bottleneck. A delivery lead who ships without scoring is a mascot. This week is the operating system. Tomorrow's charter is the first slice. Do both. The playbook that never produces a charter is a blog. When you say no, give a door that is real, or people will go around you to a vendor POC and you will meet them on day 19 with a leaked production.
Worked case · stay here ~20 minutes
Ninety minutes, four scores, one no the HIPPO has to say out loud
Marcus asked for a chatbot of everything after the keynote. You ran a 90-minute discovery on the actual job: find the last Atlas decision. Helix is not in the room. You score in the open. The everything-bot is a 1 on feasibility. The decision-log slice is fundable.
You put ninety minutes on Priya's calendar with times in the invite, not a free conversation that becomes a product pitch. In the room: two PMs who do the work, Sam, you, Jordan optional and quiet. Out of the room: Helix, and Marcus for the first hour. Agenda: walk the current task (15), pain and frequency (10), data where it lives and class (15), what good means in their words (10), thin-slice options including non-AI (15), risks they fear (10), score and next step (15). Vendors do not belong in discovery. A vendor steers toward their object and you leave with a feature list instead of a job. You already know not to invite a systems integrator to problem-framing if you want an honest problem. Marcus tried to bring Helix anyway. You parked them in Monday's kill meeting. This hour is the work as it is done today.
You watch a PM search. Slack, Confluence, a slide deck, a WhatsApp screenshot, twenty-two minutes to find whether Bravo's checkpoint moved. She finds the status slide first — the one that sent the wrong date on Friday — and only later the decision log. That twenty-two minutes is the job. The three systems are the corpus. The definition of good is a decision with a citation in under two minutes. You cannot get that from a survey that asks whether people would use AI. You time it. You write it. That sentence becomes an eval row the same afternoon: input, expected source Atlas-Decisions page Bravo-checkpoint, task success under two minutes with a citation. Frequency: this hunt happens several times a week per PM, twelve PMs. Pain: they send the wrong date, as you already know. Value is here, not in a chatbot request. You write frequency times minutes on the board so Marcus cannot later say the pain was anecdotal.
Data and class, with Sam in the seat. Atlas-Decisions is confidential when it names counterparties; for v1 you will strip names or keep the audience at the eighty. Atlas-Runbooks internal. Atlas-Program internal. Atlas-People out. Slack and WhatsApp are not in v1; they are a different class conversation and a different ACL nightmare. Owner: Sam, backup named. Sample of twenty still in progress from last week; feasibility has to score the wait. You do not pretend the drive exists so the model workstream can report green. Junk and secrets: the 2017 SOP and the key are already findings. What good means in their words: 'I would trust it if it showed the page.' Citation requirement. 'I would use it if it did not make up owners.' Groundedness and refuse. 'Do not let it mail the customer.' Never-auto, already earned. You write the sentence they said, then the metric. That is how the golden set gets rows without you inventing them in a cafe.
Thin slice and the no-AI alternative, because if you cannot name one you did not understand the job. No-AI: a required decision-log template, a single Confluence space with a naming rule, a search improvement, a form that files the decision before the status slide exists. Some of that should ship anyway. AI slice: Q&A over Atlas-Decisions and Atlas-Runbooks, citations, refuse when empty, twelve named PMs, shadow then assist, no writes. Not a chatbot of everything. Not Slack. Not mail. Risks they fear: leak of a counterparty name, another wrong date, looking stupid in steering, looking replaceable. You do not skip the job-loss fear and call it culture later. v1 is assist. v1 is not a headcount plan. If someone is planning headcount off v1, that conversation does not hide inside this charter. Hidden headcount plans are how champions become opponents. Tell the truth about the job design in the room, while they can still hear it.
You score on the whiteboard before people stand up. Four dimensions, 1 to 5, one-line why, published rubric. Decision-log Q&A: value 4, PMs lose hours and already sent a wrong date. Feasibility 4, corpus owned, class-able, eval-able now, data-ready wait of days not months. Risk 2, internal read, citations, no writes, customer send already cut. Change 3, Priya will HITL, two champions in the room, hours still tight. Fund. Chatbot of everything: value unknown, nobody will help define good. Feasibility 1, mixed class, cannot ACL, not eval-able. Risk 5, unbounded corpus, write-tools in the keynote. Change 1, no owner, no roster. Kill. You kill anything high value that is not eval-able. Executive assistant that just gets it is a wish that will consume a year. Feasibility includes eval-able. Do not hide that inside technical complexity. Let people argue with the rubric, not with you. Easy trivia consumes the roster you need on Atlas; loud is not a score.
Marcus joins for the last fifteen. HIPPO is not always wrong. It is always unaccountable unless you make it accountable. You show feasibility 1 and risk 5 on the everything-bot. You offer the adjacent fundable slice. He still wants a line in the pack about a platform vision. You say he can override in writing: I am funding the 2 despite the rubric because of X. A quiet just do the CEO one too is how the portfolio dies. You will still lose some of these. You will lose fewer, and you will have a paper trail. Refusing a control violation is the job — restricted HR with no path stays out even if he overrides the score. Sulking about the override is not. You give a door back in: finish the data-ready checklist on a named corpus, re-score. Without a door people go around you to a vendor POC, and you meet them with a leaked production. You just did.
RAID, same template, new rows, each tied to a control you have already designed. Hallucination on a high-cost act: never-auto on send, groundedness bar, citations required; trigger eval drop or a SEV. ACL leak: split indexes, identity on retrieve, ACL rows in the set; trigger a user sees a page they cannot open in source. Cost runaway: routing, cache, daily cap with a page. HITL theatre: named roster, freeze-to-assist if coverage drops; trigger queue overflow or backup on leave. Vendor deprecation: pin, drill, successor eval. Stale corpus: Sam, recrawl, sample of twenty, last-touched filter; the model will speak a 2022 SOP with more confidence than the page ever had. Assumption, treated as false: users will cite-check. If your mitigation for hallucination is the PM will notice, you have an assumption, not a control. Dependency: Sam's hours, Dana's unlesses, labeling time. Orphan rows are theatre. Tie them.
Change is trust plus a roster, in order. Shadow two weeks earns traces and a quiet comparison. Assist four weeks earns a habit because the human still sends. Champions, office hours, a visible kill switch earn permission to keep going when it misses. Benefits tracking — time, quality, avoided rework — earns the second year. Skip the order and you ask for adoption with nothing in the bank. Three champions, time in their manager's view. Office hours twice a week for the first month, agenda: bring the bad answer. Communications say what it will refuse, not only what it will do: no HR salary, no customer email, sources or I do not know. Over-promising it knows our company recreates the magical mental model. The kill switch is a change artifact: users who know you can turn it off will try it. Users who think they are stuck until next quarter will campaign against it.
Sequence and kill, published while people still like each other. Shadow two weeks, assist four, maybe confirm on Jira later, never-auto on customer send this quarter. Shadow exit: recall and groundedness at the bar, traces from real work, a champion who will say out loud whether it would have helped. If you cannot exit shadow, you do not enter assist. Entering assist because the calendar moved is how week one becomes two public misses and a trust hole you will not fill this quarter. Assist exit to confirm needs reject reasons not dominated by wrong source, a roster that actually reviewed, a written envelope. Kill criteria in numbers: groundedness under 0.8 two weekly samples in a row, or must-refuse fail: pause. Fewer than eight of twelve named users active after three weeks of assist, without a reason you accept: pause. ACL leak or SEV on a send: kill unless waived. Daily cap three days in a week: pause. Two pauses without a fix: kill unless Marcus waives with expiry.
A portfolio of more than three active AI slices on Priya's HITL roster is a fantasy. Capacity is math. Score, sequence, park. Parking is a decision with a date: later, after Atlas is in assist and the roster has slack. On the backlog is a grave. You start a public board: funded, data-ready waiting, parked, killed. The everything-bot goes in killed with the door on the card. Helix goes in killed as a vendor path, not as a use case. Decision-log Q&A goes in funded. A second idea from a PM — draft status email — goes in parked until assist is honest, because Friday is fresh. A board that only shows funded work is how HIPPO assumes everything lives. Saying no without theatre is a short written no plus a door. A long slide about strategy is theatre. People need the door. You already run intake this way; do not let AI ideas skip it because they are exciting.
Your credibility is the rubric plus one shipped slice. A delivery lead who scores everything to death and ships nothing is a bottleneck. A delivery lead who ships without scoring is a mascot. This week is the operating system. Tomorrow's charter is the first slice. Do both. The playbook that never produces a charter is a blog. You end the ninety minutes with a pack, not with this was great let's keep talking. Job-to-be-done, corpus note, first eval definition, four-dimension score, no-AI alternative, RAID rows, sequence, kill, next step: fund the thin slice. Meetings that end in exploration are how twelve POCs start. You are the person who puts a number on the board before people stand up. When a hiring manager says walk me through how you would intake AI use cases, this hour is the answer. Draw it. Do not talk about frameworks with names. Talk about the meeting, the table, and the death of a bad idea.
You send the pack the same afternoon, owners on every heading. Marcus gets the override sentence in an email so if he uses it, it is written. Sam gets the data-ready dates. Priya gets champion names and office hours to put in her team's calendar. Dana gets the risk score and the never-auto line. Jordan gets non-goals so he does not start a platform over the weekend. You keep the public board in the same folder as the handling note, the golden set, the SEV page, the SLO sheet, and the Helix kill. Tomorrow you assemble those into a charter four readers can mark up. Today you did the actual job: discovery that did not end in they want a chatbot. You will be tempted, still, to skip the no-AI alternative because it feels disloyal to the brief. You wrote it anyway. If a form and a naming rule would do half the job, the honest delivery lead says so. That sentence is worth more in an interview than a clever agent diagram.
Diagram
What you put on the board before people stood up
| Low feasibility | High feasibility | |
|---|---|---|
| High value | Everything-bot: data-ready not a model; door is a classed corpus and a re-score. | Decision-log Q&A: fund. Shadow then assist. Never-auto on send. |
| Low value | Helix-as-platform: kill. Not a use case. | Status-email draft: park until Atlas assist is honest. Easy trivia can wait. |
Fund the top-right. Send high-value low-feasibility to data-ready. Kill the everything-bot in public, with a door.
Practice
Discovery pack on a use case you know
55 minutesPick a painful, repetitive knowledge or drafting task from a program you have actually run.
- Write the 90-minute agenda with times.
- Job-to-be-done, corpus, class, no-AI alternative.
- Score the four dimensions with a one-line why each.
- RAID: 5 AI-specific rows.
- Sequence: shadow / assist / confirm with kill criteria.
Done looks like: A pack you could run Monday. This is the spine of tomorrow's charter.
Check yourself
Attempt in your notes first. Reveal is for after, not during.
What must discovery produce besides 'they want a chatbot'?
Which use cases do you kill even if valuable?
What is the default sequence?
Name four AI-specific RAID rows a DM should not omit.
How do you handle HIPPO without theatre?
What makes a kill criterion usable?
Terms from this day
- Use-case scoring
- A published rubric (value, feasibility, risk, change) used to fund or kill ideas.
- HIPPO
- Highest paid person's opinion — a governance failure mode.
- Shadow → assist → confirm
- The default intensity ramp for AI features.
- No-AI alternative
- The process, search, or form you'd ship if models did not exist. A test of whether you understood the job.
- Kill criteria
- Pre-agreed numeric conditions under which you pause or stop, published before you start.
- Job-to-be-done
- The user's actual task in one sentence, not the requested object ('a chatbot').
Your notes for day 20
Saved on this device. Use this as the start of the artifact.