Skip to main content
DoneThat

AI Adoption GuideManufacturingInspect

Inspection Escape Root Cause Correlation

ML links escape events to upstream variables such as tool wear, material lot, line speed, and operator to pinpoint defect origin and close the loop to production.

Manufacturing processPlanSourceMakeInspectPackShipServiceReturn

By Don, DoneThat’s AI coach · updated

Overview

An inspection escape is a defect that left the plant, or at least left the station that was supposed to stop it. By the time quality sees the complaint, the warranty return, or the customer containment, the line has already changed. Tools have been indexed. Lots have been consumed. Operators have rotated. The job of this use case is not to guess a root cause from a narrative. It is to reconstruct which upstream conditions were present when the escaped units were made, then test which of those conditions actually distinguish escapes from the rest of the population.

Quality still owns the corrective action. Correlation ranks candidate origins and shows the evidence. It does not issue the 8D, change the control plan, or release the line.

What the model is answering

The practical question is: given this escape event (or this cluster of events), which process variables were over-represented on the affected serials, and which were not? Typical candidates are tool wear or insert life, material lot or heat, line speed or cycle time, station, fixture, shift, and operator. The output is a ranked list of factors with effect size, sample counts, and a confidence note, not a single blamed person or machine.

That distinction matters in a quality review. A strong association with a worn insert on a specific spindle is a production input. A strong association with one operator on one shift is also a production input, but it usually points at training, work instruction, or a station that is hard to run, not at an individual as the defect origin. Treat people as a proxy for conditions you have not instrumented yet.

The loop back to production is closed only when the ranked factor can be acted on at the source: change a tool-life limit, quarantine a lot, slow a cell, or add an in-process check. If the ranking cannot name a controllable input, you still have an escape. You do not yet have a cause.

Reconstruct the escaped units before you correlate

Start from the escape, not from the SPC chart. Identify every serial, batch, or shipment that failed at the customer or at a downstream audit. Map each unit to a production timestamp, work order, routing, and station sequence. Then join the process history that was true at that moment: tool counters, lot genealogy, programmed and actual speed, operator badge, and any in-station measurements that were recorded even if they passed.

Genealogy is the usual failure point. If a lot is issued to a work order but not to a serial, you can correlate at lot level only. If tool life is reset without a reason code, wear is noise. If two operators share a login, operator drops out of the model. Fix those joins before you interpret coefficients.

Compare escaped units to a control set from the same part family, same routing, and a time window wide enough to include the same tools and lots, but not so wide that the process has been rebuilt. A control set from a different plant or a different revision will light up every difference that is not the escape.

When the escape sample is too small, return an empty correlation and say so. A handful of field failures against thousands of good units will overfit whatever happened to be present that week. Do not publish a ranked cause from n that cannot survive leave-one-out. Hold the event, keep collecting matched serials, and wait for a second independent escape or a larger containment population. An empty result is more useful than a confident story built on three units.

How to read a correlation in a quality review

Ask three questions of every ranked factor: how many escaped units carried it, how common it was among good units, and whether it is a cause or a coincident.

Counts come first. A factor that appears on every escaped serial and on a small share of good serials is a candidate. A factor that appears on every escaped serial because it appears on every serial that week is not. Report both the escape count and the baseline rate. Quality reviewers will dismiss a ranking that hides the denominator.

Effect size is next. Rank by a measure that stays stable when classes are imbalanced, and show intervals or resampling spread so a single outlier lot does not look like a process law. If two factors move together (same lot always run on the same worn tool), say they are confounded. Do not pick a winner to make the slide cleaner. Split the next run so lot and tool are no longer locked, or pull a historical window where they already diverged.

Then separate origin from detection. A correlation with final-inspection operator can mean the defect was made earlier and only some inspectors miss it. A correlation with a machining tool more often means the geometry was produced there. Use the routing: if the defect mode is a feature cut at operation 30, do not close the investigation on a pack-out scan at operation 90 even if that scan is statistically loud.

Vendors already sit on pieces of this join. InfinityQS holds the SPC and often the process variables around the inspection. SAP QM holds usage decisions, inspection lots, and the quality notification that starts containment. Qualityze holds the nonconformance and CAPA workflow once quality has a named issue. None of those systems, by itself, reconstructs serial-level tool wear against a customer escape. The correlation layer sits across them. It should write its ranked factors and evidence back into the notification or NCR so the CAPA record shows what was tested, not only what was concluded.

Quality issues the action; production changes the process

When a factor clears the sample and confounding checks, quality opens or updates the corrective action. The model can suggest a likely origin and a recommended containment (hold remaining pieces from lot X, index tool Y, 100% sort for characteristic Z). A person still decides whether that containment is proportionate, whether suppliers are notified, and whether the control plan changes.

Production then executes the process change: tool-life standard, lot break, speed limit, fixture repair, or extra check. Feed the result back. If escapes stop after the change and the factor’s association collapses on new units, the loop is closed. If they continue, the ranking was incomplete. Re-run on the new escapes instead of defending the first cause.

Do not auto-close CAPA from a correlation score. Audit and customer evidence still govern disposition. Keep the model’s output as an attachment: candidate list, sample sizes, confounded pairs, and the empty-correlation cases you declined to act on. That record is what a later 8D or a customer quality engineer will ask for.

Failure modes quality should refuse

Refuse a ranking when escaped serials cannot be tied to a timestamp or genealogy. Refuse it when the control set is a different product. Refuse it when the only significant factor is a shift or a person and no process variable was in the feature set. In that last case, instrument the station before you write a training CAPA.

Refuse real-time line stops from this model. Escape correlation is a forensic join. It runs after you have identified failed units. In-process anomaly detection and first-pass yield prediction are the tools for interrupting a shift while it is still running. Mixing those jobs produces either delayed stops or false stops.

Watch for lot-level illusions. One bad heat can dominate a month of escapes and look like a plant-wide process problem. Stratify by lot first. If one lot explains the cluster, the origin is incoming or a supplier process, and the next page you want is scrap-risk on incoming material, not a line-speed debate.

Is this worth automating for you?

Whether this pays back depends on how much time it takes your team today. Most teams estimate that from memory, and the estimate is usually wrong in one direction or the other.

DoneThat reconstructs where the time actually went, with no timers to forget, so you can measure the baseline before committing to a project and check the gain afterward.

Measure the baseline first