Use this as a written practice worksheet
The step-by-step editor needs JavaScript. You can still use this complete exercise: write an attempt and explanation, compare it with the rubric, revise, then answer the fresh challenge in your own document.
Which opportunity should the team test first?
A fictional support team wants an AI assistant. Interviews suggest routine access questions involve repeated source lookup, while exception requests require an owner decision. No baseline has been measured yet.
- The team can rehearse synthetic cases with five consenting operators.
- AI may suggest text, but the operator approves every response.
- A successful demo would not establish customer-wide time savings.
Name the user, desired outcome, comparison, observations, and continue/revise/stop rule.
What assumption could make this opportunity fail, and what observation would change your decision?
- Automate all access requests: Measure how many replies the model can produce.
Scripted guidance for this choice
Reply volume skips the user outcome and combines routine lookup with owner decisions. Narrow the opportunity and measure time through checking and correction, with a quality stop rule.
- Test source-linked help on routine questions: Compare total effort with manual work and observe unsupported claims.
Scripted guidance for this choice
This is a useful bounded experiment. Your brief still needs the user, baseline method, matched cases, the assumption being tested, and the rule for continuing, revising, or stopping.
- Ask whether operators like AI: Use positive survey responses as the go-ahead.
Scripted guidance for this choice
Stated preference does not tell you whether people can use the result or save effort. Observe operators doing matched tasks; use their explanations to understand the behavior you see.
Review criteria
- A user outcome: Does the brief measure an accepted reply, including review and correction, rather than model speed?
Measure total operator minutes from opening the request through an accepted reply, plus critical errors.
- A fair small test: Are there matched manual and assisted cases, a baseline method, and an explicit untested assumption?
Counterbalance matched synthetic cases across five operators; test whether source links reduce checking effort.
- A decision rule: Do observations lead to a clear continue, revise, or stop decision without overstating a small study?
Revise if checking erases the gain; stop on an accepted unsupported claim; a promising result justifies a bounded next test.
Revise your attempt
Sharpen the user outcome and the decision rule. Make it clear what you will observe, what AI will do, and which decision stays with the team.
The average improves, but exceptions get worse.
In a new synthetic rehearsal, routine questions are faster with assistance, but exception requests take longer and an operator accepts one unsupported policy claim.
- The small rehearsal is not representative enough for an impact claim.
- The original stop rule treated accepted unsupported claims as critical.
What will you report, stop or narrow, investigate, and test next?
Explain which original assumption changed.
- Honor the stop rule and inspect the failure: Keep the issue visible, separate case types, repair, and rerun before expanding.
Scripted guidance for this choice
The exception matters even if the average looks promising. Explain whether a narrower routine-only workflow can be tested safely after the failure cause and boundaries are addressed.
- Report the average gain and expand: Treat the exception as a minor edge case.
Scripted guidance for this choice
The average hides both segment harm and the agreed critical stop condition. A product decision needs those failures visible, not averaged away.
- Abandon every form of assistance: Assume the entire opportunity is disproven.
Scripted guidance for this choice
The result blocks expansion but does not explain every possible bounded design. Inspect the failure and segment the work before deciding what, if anything, deserves another test.
Give the revised brief to an engineer and an operator. Ask each to describe the test they would run without additional explanation.