Expressive ribbon artwork, separate from the recorded waveform below.
Sunview / Call excerpt
Shipment reference muted / Original timing
0:00 / 0:19
Call transcript 0:19 / Redacted excerpt
Agent: Thanks for calling Sunview Logistics. How can I help you today?
Caller: Hi, can I check up on a load status?
Agent: Yeah, I can do that. What's your load or reference number?
Caller: [Load reference muted.]
Agent: Alright, one moment. Right, [load reference muted] is delivered.
Automatic transcript, edited to mark redactions. 19-second excerpt. Shipment reference muted; original pace and pauses preserved. The source is a 48 kHz Opus recording.
Measured, not researched.Built, questioned, tested.
From the workbenchResearch note / 01
A better model. Or a better question?
I tested speech models on the situations our voice agents actually encounter. A fast transcription isn't much help if it misses the thing a caller needs.
40 synthesized domain clips; a 7-clip name subset for keyterms
Public harness Release pending
Latency / word error rate
Parakeet TDT 0.6B v3170 ms / 0.047 WER
Three reported baselines; one keyterm projection. Local and hosted runtimes differ.
Method & limits
Historical reported summary from open-weight-stt-findings.md, not a fresh run. WER is a fraction; latency is the reported median in milliseconds. Whisper and Parakeet ran locally on an M4 GPU; Deepgram used a hosted API, so timing includes different compute and network conditions, not live-turn latency. The keyterm row is the same Deepgram system: full-corpus WER is projected from seven name clips, and 204 ms is inherited from the baseline, not remeasured. The separate 240-row probe pools six audio conditions and is not this 8 kHz summary.
A Chrome extension that checks article claims against sources and places the results beside the text.
I chose Perplexity because I preferred its search at the time. Bring your own key kept me from taking on ongoing API bills, and the extension did not need a backend of my own.
I'm DongYeop Lee, or DY. I spent Summer 2026 at Hemut working on voice agents and the backend behind them. That meant moving between conversations, code, and the tests that were supposed to tell me whether any of it worked.
Seeing issues in Converse sent me back to my own benchmark. I had already reached a conclusion, but the scorer needed another look. I care about that part of building just as much as getting the first version running.
This site is a project, too. I'm building mine first so my designer partner and I can learn from it before we build hers. Veracity and EcoPrompt are here for the same reason: I wanted to make the idea real and see what happened.
“I called my own agent. I got tired of waiting.”Why I went after the latency.
Recent focus
Voice agents Backend infrastructure AI evaluations
Also built
NLP at IDX Exchange A browser extension A hackathon with a designer