Skip to main content
OCT 2026 Latest guide updates. Latest: Mistral Large 4 AI FinOps Lean & AI Changelog →
← Resource library

Lean × AI · reading path

Follow the work to the user

AI can make one step faster while the request waits somewhere else. Map the full path before calling that a productivity gain.

The operating question

What reaches the user sooner?

A draft, a pull request and a support reply are intermediate outputs. The user sees a working feature or a resolved problem. When AI increases the flow into review, testing or approval, the queue at that stage may grow even if creation time falls.

The queue shift is a possibility, not a law. Measure both the local cycle and what happens after it; the field study and team accounts below show why one number cannot settle the question.

Lean gives us a way to trace that request. Toyota’s principles offer a useful analogy: make what is needed, when it is needed, and stop defects before they travel downstream. They do not supply a software-team performance forecast.

A five-stage map

Trace one request end to end

Choose a real request type. Record when it enters and leaves each stage, including revisions and rejected work. Compare equivalent requests over equivalent windows.

Factory analogy: a demand card enters an AI agent cell. A quality gate with an andon lamp can stop and return faulty work for rework. Accepted work is released before the user outcome is checked.
Read the Lean signals. The pull signal evokes kanban: the next stage requests what it needs. The lamp evokes andon; jidoka also requires stopping and correcting affected work. The shipping dock is not the user outcome. Measure the agent-cell cycle and the full path. Enlarge factory diagram ↗

Swipe sideways to explore the diagram, or use the enlarge button.

The same request, in software

Record where time and corrections accumulate

  1. 01

    Demand

    Which user problem is worth solving?

    Requests clarified or declined before work starts
  2. 02

    Creation

    What does AI make faster?

    Active work time and items submitted
  3. 03

    Verification

    Can someone establish that the result is sound?

    Waiting time, reviewer effort and corrections
  4. 04

    Delivery

    Did the work reach its intended user?

    Time to release or response; failures
  5. 05

    Outcome

    Did it solve the original problem?

    Adoption or resolution for eligible users

The next constraint

When drafts arrive faster than decisions

The queue between creation and acceptance is the part to inspect next. Count items waiting, how long they wait, and the effort spent checking and correcting them. The drawing shows one possible failure mode, not a forecast.

Queue balance · illustrative counts

What is still waiting for verification?

Factory analogy: an AI agent cell sends draft cards to a waiting rack before a quality gate. Accepted work exits. If arrivals exceed exits, waiting work grows. The cards show no measured counts.

Swipe sideways to explore the diagram, or use the enlarge button.

Read the Lean signals. The shelf is work in progress (WIP), not automatically waste. If measured flow shows unused output, uneven arrivals or reviewer overload, investigate muda, mura and muri. The blank cards show no measured volumes. Enlarge queue diagram ↗

Example, not field data. This count assumes no cancellations or other exits. A correction remains part of the same item. Also track the oldest wait, review effort and defects; counts alone cannot establish user value. Define the queue measures.

Decision rule

If creation gets faster but verification, correction or waiting gets worse, adjust the amount of work admitted or the verification path before scaling generation.

Lean software development

The terms behind the factory analogy

Mary and Tom Poppendieck’s software adaptation calls for building quality in and optimizing the whole value stream. Here, time to delivery and the later user outcome remain separate measures. These terms name possible controls to examine in a real workflow.

Just-in-Time · kanban
A downstream need sends a pull signal upstream. A proposed work-in-progress limit would admit another agent draft only when verification has capacity; a board alone does not create pull. Toyota · Kanban Guide
Jidoka · andon
Jidoka detects an abnormality and stops affected work; an andon makes it visible. In software, early tests and evidence checks build quality in; final review cannot replace them. Toyota · Poppendieck
Poka-yoke · error-proofing
A check close to creation rejects a known mistake before it reaches review. For an agent, a reproducible build or a targeted test is useful only if it catches the failure it was added to prevent. Toyota · Lean Enterprise Institute
Muda · mura · muri
Waste, unevenness and overburden. Unused drafts or avoidable rework may be muda; bursts of submissions, mura; an overloaded review path, muri. Check the actual flow before applying these labels. Lean Enterprise Institute
Kaizen
When a failure recurs, people change an upstream test, agent instruction or intake rule, then check whether recurrence falls. An alert alone is only a signal. Toyota

Practice, with its limits

What companies have actually described

Lean can shape how teams build and verify AI-assisted work. AI can also help people practise Lean problem solving. These sources document implementations and practitioner experience, but they do not establish a transferable end-to-end productivity gain.

One field study, two engineering accounts

The study reports a task-cycle comparison. Qonto and Theodo describe changes to their own work. Read each finding with its limit.

Software field study

Globo: a shorter task cycle

In a six-week study, Jira tasks tagged as using GenAI had an average development cycle time 23% below the company’s historical figure.

Evidence limit: Participation and task tagging were voluntary. A historical comparison does not establish that AI caused the difference or improved release quality and user outcomes.

Read the ICSE-SEIP study ↗

Practitioner account · software migration

Qonto: migrate code through kaizen

A small team iterated from a chatbot to an AI-assisted CLI with a codemod and human review. One engineer, with limited help, migrated 8,632 lines in a two-week evaluation.

Evidence limit: The manual baseline was estimated. Lines migrated do not measure review cost, defects or customer value.

Read Qonto’s engineering account ↗

Practitioner account · agent workflow

Theodo: stop and correct the process

On the AP-HP ComPaRe EMA project, a CTO reports shipping 15 features in three weeks while reviewing strategy, tests and code, and changing agent instructions after failures.

Evidence limit: The two-month comparison was an internal estimate. User adoption and later defect rates were not published.

Read the Theodo account ↗

Related practices, different evidence

These examples illuminate an alert, a quality decision and a coaching method. They do not measure the end-to-end software flow shown above.

Lean applied to AI work

Toyota: observe the agent system

Toyota’s enterprise AI team describes LangSmith as an andon board for failed tools, broken pipelines and agent adoption across its internal platform.

Evidence limit: The case is published by LangChain, a supplier. Monitoring an anomaly does not itself prove that work stops or the cause is removed.

Read the Toyota case from LangChain ↗

Lean quality management study

Siemens: keep quality decisions accountable

A five-year practitioner study at Electronics Works Amberg examines tensions between human and AI decisions, transparent and opaque reasoning, and specification-led and data-led processes.

Evidence limit: The paper analyses adoption in one manufacturing setting. It does not measure an AI-driven speed or quality gain.

Read the Siemens study ↗

AI supporting Lean practice

LEI: critique the reasoning, not the A3 form

Art Smalley describes RootCoach as feedback on root-cause reasoning. The proposed benefit is coaching available when a human coach cannot respond.

Evidence limit: His quality comparison is personal judgment, not a blind study. LEI also warns that an AI coach can displace human authority.

Read the RootCoach account ↗

There are inspectable open-source implementations too. Andon records recurring agent defects and countermeasures; its completion hook has stated limits when the agent writes its own verification entry. RaiSE describes a Lean-inspired software workflow. These repositories show design choices, not measured improvements in user outcomes.

Evidence, with its limits

Read the source for the claim it actually supports

Principle

Toyota Production System

The primary source for Just-in-Time and jidoka. Its manufacturing outcomes do not transfer automatically to software.

Research

DORA: AI changes the shape of work

Qualitative responses from Google engineers describe time moving toward auditing and verification. Check the scope before generalizing.

The Microsoft rollout study measured about 24% more merged PRs among adopters over four months. Merged PRs remain an output proxy. Use DORA’s delivery metrics for the release portion and define a separate user outcome for the complete service.

Choose your next step

Test the idea on one kind of request

  1. Choose a request. Use the same request type before and after a change.
  2. Trace its path. Record arrival, waiting, correction, release and the later user outcome separately.
  3. Change one control. Adjust admission or verification, then compare the same signals again.

Before adding more agents

Make one agent's station testable

Marek Kalnik's Lean talk describes shared standards and project readiness as conditions for AI work. Eno Reyes argues for reliable checks on a single task before parallel agents. Use Factory's repository criteria to find a concrete gap, then test whether fixing it changes the work.

  1. Observe a failureFind a repeated setup guess, inconsistent convention, flaky check or defect that reached review.
  2. Fix the local controlA repository owner selects the failure to prevent. Make the command reproducible or add a check that catches that mistake on a representative task.
  3. Follow the requestCompare retries, tokens and reviewer corrections for similar work, then track release and the user outcome.
Example: one legacy change

A source-line check can prove that a citation exists while the agent still misreads the behavior. Before migration, have a reviewer follow the relevant call path and run characterization checks on the existing application. Move one usable unit through review and release before opening several agent lanes. If the format is fragile, exercise the real edit path and check the resulting bytes, including alternate writes.

These are proposed controls, not results from a deployed migration. Record false blocks, rework and reviewer time alongside accepted changes.

Evidence limit: A readiness score describes conditions in the repository. It does not measure agent productivity, shorter delivery or value to the user. Factory's five-level model is its own rubric; scores from another radar are not directly comparable.

From my open-source work

Know what each instrument can see

These tools expose parts of AI-assisted work, but none connects a request to acceptance, release and the user outcome. Use their signals in a local value-stream trace, not as a productivity score.

ctxharness ↗
Checks agent instruction files for stale versions, broken paths and missing scripts. A sounder starting context does not prove the resulting change works.
cc-skill-usage ↗
Counts actual Skill tool invocations in Claude Code transcripts. An invocation is evidence of use, not of a correct decision or accepted result.
ccboard ↗
Shows session activity, token use and costs. Those operational signals need to be joined to verification effort, delivery and product outcomes elsewhere.

For the measurement evidence, read my AI productivity review. For a human check at the quality gate, the UVAL protocol tests whether an engineer can explain the code they accept; it does not replace team-level review or a user outcome.