Skip to content
All training paths

Operator training · agenda published · pilot cohorts forming

AI Workflow Evaluation: from promising demo to operating decision

A practical path for finding one measurable workflow, testing it against a baseline, and deciding whether it deserves budget, integration, and governance.

Who it is for

Bring a live decision.

  • Operations and product leaders choosing where to deploy AI
  • Finance and transformation teams reviewing AI business cases
  • Team leads responsible for adoption, quality, or measurable ROI
  • Consultants and internal enablement teams designing pilots

Format

Four 60-minute modules. Available as a facilitated team study or two half-day working sessions.

Before the session

Bring one recurring queue or workflow. No coding experience is required.

Launch status

The agenda and completion standard are published now. Facilitation notes and exercises are being refined with early cohorts, so we will confirm fit and format before proposing a session.

Course agenda

Four modules. Four working artifacts.

01

60 minutes

Find work worth measuring

Separate a real operating queue from a broad “use AI” ambition.

  • Queue volume and wait time
  • Named owners and reviewers
  • Reversibility and cost of failure
  • Use-case rejection rules

Workshop: Score three candidate processes and select one clearly limited job for the course.

Take-home artifact: Workflow candidate scorecard

02

60 minutes

Build the baseline before the pilot

Define the current cost, cycle time, quality, and retry rate before introducing a model.

  • Baseline sampling
  • Loaded labor cost
  • Quality definitions
  • Hidden review and correction work

Workshop: Measure a small sample of the current workflow and document its confidence limits.

Take-home artifact: Baseline evidence sheet

03

60 minutes

Test the failure modes, not the demo

Run normal, edge, and adversarial cases with a documented human review step.

  • Representative test sets
  • Accepted-output rate
  • Severe error classes
  • Escalation and stop conditions

Workshop: Create a 12-case evaluation set and a rubric another reviewer can use.

Take-home artifact: Evaluation set and scoring rubric

04

60 minutes

Make the buy, redesign, or stop decision

Translate the test into a decision that includes seats, integration, review, and failure cost.

  • Cost per accepted output
  • Sensitivity ranges
  • Pilot-to-production gaps
  • 30-day operating review

Workshop: Write a one-page recommendation with a funding ceiling and stopping rule.

Take-home artifact: Evidence-backed adoption memo

Capstone and completion

One workflow decision packet

The capstone combines the scorecard, baseline, evaluation results, economics, owner, review design, and stop rule for one real workflow. A polished demo without comparative evidence does not pass.

Completion standard: Completion means submitting a reviewable decision packet for one real workflow. It is a course completion record, not an accredited certification.

Frameworks used in the course

Pilot cohort

Shape the course around work you already own.

Send the team size, use case, and decision in front of you. Gary will reply by email with fit and possible next steps. This does not reserve a date or collect payment.

This opens a pre-filled email to Gary Stanton. It does not book a time automatically; you will receive a personal follow-up.