Worked example · anonymised and illustrative · the shape of an engagement, not a past result
Route assistant: a recommendation layer dispatchers actually accept, measured depot by depot.
Dispatchers planning a thousand-plus routes a day against a TMS that cannot explain its own suggestions. The question is not whether a model can plan routes but whether dispatchers will take its advice, so acceptance rate is the number the engagement would be paid on.
Measurement plan · agreed before build, measured weekly
| Metric | Baseline | Target | How measured |
|---|---|---|---|
| Depots live | 0 | 1 by week 6; all pilot depots by week 10 | Depots with assistant enabled for all dispatchers |
| Recommendation acceptance | Measured in week 6 | Agreed at gate 2; rising at every later gate | Accepted ÷ shown, 7-day trailing |
| Empty-leg kilometres | Prior-year same week | Agreed at gate 2 | vs same depot, same week, prior year |
| Inference cost per recommendation | Measured from first live day | Under the SOW ceiling by week 10 | Token spend ÷ recommendations, incl. retries |
A recommendation the dispatcher ignores is a cost, not an outcome. That is why acceptance, not model accuracy, is the paid-on number.
The situation this is written for
A transport management system whose route plans dispatchers override about half the time, usually for reasons the system cannot see: a customer who only unloads before nine, a driver who knows a yard.
How we would run it
Weeks one and two are discovery: a lineage map of dispatch data, a costed plan, and a go/no-go on the findings. Weeks three and four build the eval set from historical routes and the overrides against them, and publish baseline metrics and the cost model: gate 2. The first production slice goes live in one depot in week six, with the cost per recommendation measured from the first day.
The assistant explains every recommendation in the dispatcher’s terms and records the override reason when it is rejected. Those reasons feed the weekly eval review. We expect the largest gains in acceptance to come from changing what is shown and how it is explained, not from changing the model, and the plan budgets for that.
What you would be left with
Drift alerts, a dispatcher feedback loop the operations lead owns, runbooks, and an exit review delivered in writing in week twelve. Your engineers run on-call unaided for the final two weeks.
Start with a two-week discovery sprint.
Same people scope and build. Numbers agreed up front, measured weekly, published at exit.