Envion Software
CS-065Machine Learning ConsultingPharmaceutical Manufacturing (NDA)

Visual Inspection: The Model Wasn’t the Hard Part

A pharmaceutical packaging manufacturer running manual visual inspection on blister packs had 240,000 labelled images and a vendor quote for a deep learning inspection system. Envion’s feasibility read: the labels were pack-level dispositions, not defect locations; the defect rate was 0.4% — about 960 positive examples across seven classes, two with fewer than 30; and the real blocker was acquisition, not modelling — specular highlights, skylight-driven lighting variation and focus drift meant the photographs didn’t reliably contain the signal (their own senior inspectors, tested blind, disagreed with the recorded disposition on 23% of marginal cases). The verdict: yes, in two stages — fix the cameras first, then automate the three high-signal classes as an assistant, not an auto-reject.

Visual Inspection: The Model Wasn’t the Hard Part
01

The challenge

The client ran manual visual inspection on blister packaging: seal integrity, tablet presence, print legibility, foil damage — two inspectors per line per shift, with escaped defects reaching customers at a rate their quality director described as "low but not zero, and not zero is the problem in pharma."

They had 240,000 labelled inspection images from the past two years and a vendor quote for a deep learning inspection system. They wanted an independent read before signing.

The board-level question was "can machine learning do this?" — which is almost never the useful question. ML can do a remarkable number of things badly enough to be worthless; the interesting part is the gap between "technically possible" and "worth building."

02

Decision path

The assessment answered five things: can the model learn this, does the data actually support it, how would we know if it worked, what would it take to run in production, and does the expected value survive contact with the base rate.

The images were labelled, but not the way anyone assumed: labels came from the QA disposition record — pass or fail at pack level — not from anyone marking a defect location, and in about 8% of cases the failure wasn't even visible in the photograph. Defect rate was 0.4%: around 960 positive examples spread across seven defect classes, two with fewer than 30 examples. You are not training a reliable detector for a class with 30 examples, and no amount of augmentation changes that — augmentation multiplies what you have; it doesn't add information about failure modes you've never photographed.

The decisive test: archived images given blind to two of the client's own senior inspectors. They disagreed with the recorded disposition on 23% of marginal cases and with each other on 17%. If experienced humans can't reliably call it from the photograph, the photograph doesn't contain the signal — a model trained on that data learns the noise and reports it confidently.

03

Envion contribution

The finding that mattered took two days on site rather than two weeks with the dataset: the real blocker was acquisition, not modelling. Fixed-exposure cameras mounted at an angle that put a specular highlight across a third of every foil surface, lighting varying with a skylight above line 2, focus drifting between maintenance intervals.

Prototyping on the three classes with adequate examples and controllable appearance — tablet presence, gross seal deformation, print absence — worked well even on the poor images, because the signal is large. Fine print legibility and hairline foil perforation did not, and would not have with a bigger model.

Envion explicitly recommended against attempting all seven classes at launch — which is what the vendor had scoped — and against any auto-reject decision in year one.

04

Delivery

The verdict: feasible, in two stages, and worth doing — but the vendor's proposal would have failed, because it assumed the data problem was solved.

Stage one, before any ML spend: fix image acquisition — telecentric lighting to kill the specular highlight, enclosed illumination so the skylight stops mattering, fixed focus with a maintenance check, and a second camera angle for seal inspection. About €60K of hardware and four weeks. Inspectors also began capturing defect-region annotations — thirty seconds per rejected pack — so that in six months the client would have localized labels rather than pack-level ones.

Stage two: train on the three high-signal classes, deploy as an assistant that flags for human confirmation rather than auto-rejecting, and revisit the remaining classes once improved capture had accumulated enough examples. Production requirements: inline inference at 4 packs/second per line on edge hardware — latency and network reliability on a production floor make cloud non-negotiable in the wrong direction; model versioning tied to the batch record, because in a GMP environment you must be able to say which model version inspected which batch; a defined revalidation trigger on model update; and a drift monitor on the input images, not just predictions — if a camera degrades, the images shift before the accuracy does.

05

Outcome and evidence

Fourteen months after stage one, escaped defects fell from 340 to 41 per million packs, three classes were automated at launch and six by month twelve, human inspection load dropped to ~18% of packs (flagged only), false reject rate fell from 2.1% to 0.6%, and ~11,000 localized defect examples had been collected — the sole reason the sixth class became trainable at month twelve. The seventh still isn't, and the client was told at the time it probably never would be from imagery.

The advice that generalizes: run the blind human test before you talk to a single vendor — it costs an afternoon, and if your experts can't agree from the image, no model will fix it, you'll just be paying to have the disagreement automated. And count positive examples per class, not in total: "240,000 images" was really 960, split seven ways. That arithmetic is the entire feasibility question for most defect detection projects, and it takes ten minutes.

Results — 14 months after stage one
MetricBeforeAfter
Escaped defects (per million packs)34041
Classes automated03 at launch, 6 by month 12
Human inspection load100% of packs~18% (flagged only)
False reject rate2.1% (manual)0.6%
Inspectors per line per shift21
Localized defect examples collected0~11,000

Client feedback

What the client says about this engagement

Head of Manufacturing Technology

“We had a quote for a deep learning system and Stanislav told us to spend sixty thousand euros on lighting first. That was not the meeting I expected. The part that convinced our quality director was when he took our own inspectors, showed them our own images blind, and they disagreed with each other on a fifth of them. You cannot argue with that.

He also told us to start annotating defect locations by hand, which felt like busywork for about six months and is the only reason we automated three more classes in year one.”

Head of Manufacturing Technology · Pharmaceutical packaging manufacturer (NDA, anonymized)

From the engagement lead

What I’d tell anyone considering this

V. Stanislav

“If you're looking at computer vision for inspection, run the blind human test before you talk to a single vendor. Take your archived images, strip the labels, give them to two experienced inspectors, and measure agreement. It costs you an afternoon. If your experts can't agree from the image, your data has an acquisition problem and no model will fix it — you'll just be paying to have the disagreement automated.

The other thing: count your positive examples per class, not in total. "240,000 images" sounds like plenty and mine was really 960, split seven ways. That arithmetic is the entire feasibility question for most defect detection projects, and it takes ten minutes.”

V. Stanislav · Senior ML Engineer at Envion Software

Evidence gate. This page publishes only what Envion's project records and client disclosure permissions support. Outcomes are added once verified against a baseline, a measurement period, and an approved source.

FAQ

Questions about this case

Facing a similar challenge?

Considering computer vision for inspection? Run the blind human test before you talk to a single vendor — or let Envion run the feasibility read for you.

Discuss a Similar Challenge

Executive Technology Leadership

Support for high-stakes product and AI decisions

Bring senior technology leadership into the business when the roadmap is unclear, delivery is at risk, an AI initiative needs stronger ownership, or the company needs an experienced technical voice before hiring a permanent CTO.

Discuss Interim CTO Support

Core responsibilities

  • Align product and technology priorities with business goals and measurable outcomes.
  • Review architecture, delivery risks, data foundations, security needs, and AI readiness.
  • Lead internal teams and external partners through a practical execution plan.
  • Clarify team structure, ownership, decision rights, and delivery cadence.
  • Support investor, board, partner, and due-diligence conversations with credible technical judgment.

New experience

Prompt-to-Page — try it right here

Describe the landing page you want, in your own words. We turn it into a finished page and email you a private link in 5–10 minutes — no briefs, no calls, $0 to see the result.

  1. Describe what you want to create.
  2. We structure, write, and compose the page.
  3. You receive a private link when it is ready.

Start with a sentence — the interactive builder takes it from there.

Generate My Page

Safe, respectful content only. No obligation.

Start here

Discuss a Similar Challenge

Share your current state, constraints, and desired outcome — a senior specialist will reply with a concrete next step.

Prefer a direct channel?