← All work
Case study · in production Machine learning & AI

Who said what, on messy real-world audio

Voiceprint diarization as a service

ML infrastructure1–2 engineersVerified on GPU Cloud Run

The problem

Off-the-shelf transcription merges speakers on noisy walk-and-talk audio — and every downstream AI judgment inherits the error. When four people talk, two get merged into one.

The system

A speaker-embedding pipeline that matches transcript words to voiceprints, splits wrongly-merged turns, and abstains when attribution is genuinely ambiguous — because a wrong speaker label is worse than no label.

How it's built

Delivery

Research-to-production in short cycles: offline eval on labeled tours first, then a live shadow lane, then default-on.

Results

Client anonymized by industry. Detailed numbers and references available on a call.

Want a system like this?