The short answer

Generative AI pays back when it is scoped to a single measured job inside the product, with a baseline recorded before the first prompt is written.

Abstract

Most product teams are being asked to add generative AI without a clear definition of what it should do or how success will be judged. This paper sets out the operating model we use with enterprise product teams: pick the workflow before the model, instrument the baseline before the build, and ship an assisted version of one real job rather than a feature bolted onto everything.

Summary

Generative AI arrives in most product organisations as a mandate rather than a problem statement. The board wants AI in the product, the roadmap grows an epic, and six months later the feature exists but nobody can say what it changed.

The teams that get value do something narrower. They choose one job inside the product that has high volume and low judgement, measure how long it takes and how often it goes wrong today, and then ship an assisted version of exactly that job. Everything else, model choice, retrieval design, prompt structure, follows from the job.

This paper documents the operating model we use on forward deployed engagements: a selection matrix for choosing the first workflow, the instrumentation to put in place before any build starts, a reference architecture that keeps model routing swappable, and an evaluation harness that runs in CI so a model upgrade is a config change rather than a rewrite.

It closes with the unit economics, cost and latency per task at realistic volumes, and a twelve-week plan that takes a team from a scoped problem to an assisted feature in front of real users.

Headline findings

1

Workflow per release, the single strongest predictor of a shipped AI feature

6-10 wks

From scoped problem to an assisted version running with real users

3x

Faster iteration once evaluations run in CI instead of by hand

0

Model lock-in when routing sits behind a swappable interface

What is inside

  1. 01

    Why most AI features stall

    The failure is scope, not capability. Six symptoms to check for in your own roadmap.

  2. 02

    Pick the job, not the model

    A selection matrix for the workflow with the highest ratio of volume to judgement.

  3. 03

    Baseline before build

    What to instrument in the two weeks before any prompt is written.

  4. 04

    Reference architecture

    Retrieval, routing, evaluation and guardrails, with the boundary each one owns.

  5. 05

    Evaluation that survives a model swap

    Golden sets, regression gates and how to keep them cheap.

  6. 06

    Cost, latency and the honest maths

    Unit economics per task and where they break.

  7. 07

    Rollout and change management

    Assist, then review, then autonomy, gated on measured error rate.

  8. 08

    Twelve-week plan

    A week-by-week schedule you can lift into your own planning.

Who it is for

  • Heads of product adding AI to an existing enterprise product
  • CTOs deciding build, buy or wait
  • Engineering leads who have to maintain the thing after launch

Full report

Read the complete 24-page paper.

The summary above is the argument. The PDF carries the data, the architecture diagrams and the plan.

One email, the report. No sequence. We use the company field to send the version closest to your operation, Product.

Written by

Utah Tech Labs

Product engineering. Forward deployed engineers at Utah Tech Labs, writing from engagements that reached production.

Meet the team >