Quality Assurance Labs
AI Apps & Integration

Computer Vision in Production — The Testing Checklist

Senior AI Engineer7 min readPublished Updated

A CV model that hits 95% accuracy in the lab can drop to 60% in production. Here's the checklist we use to catch that before users do — lighting, edge devices, bias, latency, drift, and adversarial inputs.

Camera inspecting objects with detection frames
#computer-vision#edge-AI#model-testing#ML-QA#TensorFlow-Lite

Computer vision models are notoriously bad at generalizing. A model trained on clean data fails on messy inputs. A model that hits 95% accuracy in your test set can collapse to 60% in the field.

Here's the testing checklist we use at QA Labs before any CV model ships.

Test against real-world conditions

Lab accuracy ≠ field accuracy. Test against:

Real-world lighting (not studio)

Different camera angles

Motion blur

Occlusion (partial obstructions)

Weather conditions (for outdoor)

Different times of day

Test on edge devices

Latency and accuracy change dramatically on edge hardware:

Measure inference time on target devices

Check memory usage under load

Verify battery impact

Test model quantization (if applicable)

A model that runs fine on GPU servers might time out on a Raspberry Pi.

Test for bias

CV models often perform worse on underrepresented groups:

Measure accuracy across demographics

Check for false positives/negatives by group

Test with diverse datasets

Audit training data for representation

Bias testing is non-negotiable for any CV model used in decisions affecting people.

Test for model drift

Model accuracy degrades over time:

New patterns appear

Data distributions shift

Environments change

Plan for periodic retraining. Set up monitoring for accuracy drops.

Test adversarial inputs

CV models can be fooled by:

Adversarial perturbations (small pixel changes)

Out-of-distribution inputs

Trick images

For security-sensitive applications, adversarial testing is required.

What we typically find

Models trained on studio data fail on real-world images

Latency on edge devices exceeds SLA

Demographic bias in face-related models

Accuracy drops 10–20% over 6 months without retraining

Key takeaways

  • Test CV in real conditions, not just the lab
  • Edge devices expose latency and accuracy issues
  • Bias testing is non-negotiable
  • Plan for drift and retraining
  • Adversarial testing for security-sensitive deployments

Further reading

About the author

Senior AI Engineer →

Senior AI Engineer · Quality Assurance Labs

Notes from the lab.

Testing, engineering and growth — delivered to your inbox.

Need CV model QA? Book a scoping call

Let's talk →