Worked example · anonymised and illustrative · the shape of an engagement, not a past result

Applied AI delivery · European logistics operator

Route assistant: a recommendation layer dispatchers actually accept, measured depot by depot.

Dispatchers planning a thousand-plus routes a day against a TMS that cannot explain its own suggestions. The question is not whether a model can plan routes but whether dispatchers will take its advice, so acceptance rate is the number the engagement would be paid on.

Sector
Logistics
Setting
Germany and the Netherlands
Duration
12 weeks
Team
Both founders
Commercials
Fixed price, milestone-billed

Accept rate
recommendation acceptance rate: the number we would be paid on, measured weekly per depot
1 → n
depots live: one by week six, the rest only when the first holds its numbers
€ ∕ rec.
inference cost per recommendation, with a ceiling in the statement of work

Measurement plan · agreed before build, measured weekly

Metrics measured weekly
MetricBaselineTargetHow measured
Depots live01 by week 6; all pilot depots by week 10Depots with assistant enabled for all dispatchers
Recommendation acceptanceMeasured in week 6Agreed at gate 2; rising at every later gateAccepted ÷ shown, 7-day trailing
Empty-leg kilometresPrior-year same weekAgreed at gate 2vs same depot, same week, prior year
Inference cost per recommendationMeasured from first live dayUnder the SOW ceiling by week 10Token spend ÷ recommendations, incl. retries

A recommendation the dispatcher ignores is a cost, not an outcome. That is why acceptance, not model accuracy, is the paid-on number.

The situation this is written for

A transport management system whose route plans dispatchers override about half the time, usually for reasons the system cannot see: a customer who only unloads before nine, a driver who knows a yard.

How we would run it

Weeks one and two are discovery: a lineage map of dispatch data, a costed plan, and a go/no-go on the findings. Weeks three and four build the eval set from historical routes and the overrides against them, and publish baseline metrics and the cost model: gate 2. The first production slice goes live in one depot in week six, with the cost per recommendation measured from the first day.

The assistant explains every recommendation in the dispatcher’s terms and records the override reason when it is rejected. Those reasons feed the weekly eval review. We expect the largest gains in acceptance to come from changing what is shown and how it is explained, not from changing the model, and the plan budgets for that.

What you would be left with

Drift alerts, a dispatcher feedback loop the operations lead owns, runbooks, and an exit review delivered in writing in week twelve. Your engineers run on-call unaided for the final two weeks.

Start with a two-week discovery sprint.

Same people scope and build. Numbers agreed up front, measured weekly, published at exit.