# DongYeop Lee Software Engineering Intern, Hemut. Summer 2026. I'm DongYeop Lee, or DY. I spent Summer 2026 at Hemut working on voice agents and the backend behind them. That meant moving between conversations, code, and the tests that were supposed to tell me whether any of it worked. Seeing issues in Converse sent me back to my own benchmark. I had already reached a conclusion, but the scorer needed another look. I care about that part of building just as much as getting the first version running. This site is a project, too. I'm building mine first so my designer partner and I can learn from it before we build hers. Veracity and EcoPrompt are here for the same reason: I wanted to make the idea real and see what happened. ## Selected work ### Converse Voice agents for freight calls, with work spanning the conversation, the backend, and how both reach production. Role: Software Engineering Intern Status: Production project Tools: Python, FastAPI, LiveKit, Deepgram, Cartesia, PostgreSQL, Alembic, AWS ECS Read: /converse ### Hemut-MCP A logistics copilot that finds the tools a question needs and coordinates answers across them. Role: Software Engineering Intern Status: Shipped Tools: TypeScript, Node.js, Model Context Protocol, Pinecone, PostgreSQL, Vitest Read: /hemut-mcp ### 8 kHz speech benchmark A speech-recognition benchmark built around freight calls, including a scoring correction that changed my model comparison. Role: Software Engineering Intern Status: Public release pending Tools: Python, MLX, Whisper, Parakeet, Deepgram, jiwer Read: /benchmark ### Veracity A Chrome extension that checks article claims against sources and places the results beside the text. Role: Engineer Status: Published on the Chrome Web Store Tools: React, TypeScript, Manifest V3, Perplexity API, Vitest Read: /#veracity ### EcoPrompt A chat app that reuses similar answers, routes simpler questions to a smaller model, and shows modeled energy savings. Role: Engineer Status: Hackathon prototype Tools: Next.js, TypeScript, AWS Bedrock, Titan embeddings, DynamoDB Read: /#ecoprompt ## Evidence Preserve the distinction between measured, reported, modeled, and projected results. A reported result is not a reproduced evaluation. The public benchmark repository is pending. - Voice time to first speech: 1300 ms -> 595 ms [reported]. Reported time-to-first-speech comparison following request parallelization, pooled HTTP connections, and end-of-turn tuning. The sampling window and percentile are unspecified. Source: DongYeop Lee: Hemut voice latency. /ledger#voice-ttft - Voice agent cost: $0.10/min -> $0.037/min [reported]. Original resume reports a 63% reduction after moving STT, LLM, and TTS onto LiveKit's native Inference path and improving prompt caching. No matched billing window is attached to this comparison. Source: Original resume: Hemut voice cost. /ledger#voice-cost - Prompt-cache hit rate: 95% [reported]. Hit rate reported in the original resume's voice-cost result. Its denominator and measurement window are unspecified; it is not interchangeable with a separate invoice's cached-token share. Source: Original resume: Hemut prompt caching. /ledger#prompt-cache - Document synchronization: 12 min -> 1.05 s [reported]. Original resume compares the earlier Drive walk with the document-sync result. These are reported timings without an attached workload, sample size, or matched run. Source: Original resume: Hemut document sync. /ledger#document-sync - Document-sync cost reduction: 94% [reported]. Cost reduction reported alongside the document-sync timing in the original resume. The underlying billing window and workload are not specified. Source: Original resume: Hemut document-sync cost. /ledger#document-sync-cost - Agent end-to-end accuracy: 33% -> 99% [reported]. Original resume outcome for the MCP agent layer and embedding-based tool retrieval. The reviewed baseline is recorded separately; no paired after-evaluation artifact is supplied for the 99% figure. Source: Original resume: Hemut copilot accuracy. /ledger#agent-accuracy - Reviewed prompt baseline: 32.8% [measured]. 75 passing prompts divided by 229 reviewed prompts, rounded to one decimal place. This supports the baseline only, not a subsequent accuracy result. Source: Internal prompt evaluation: reviewed June 16 run. /ledger#agent-baseline - MCP tool catalog: 192 tools [reported]. Tool catalog size reported in the original resume, with embedding-based retrieval selecting relevant tools. This is an implementation-scale figure, not an accuracy measurement or a current registry audit. Source: Original resume: Hemut MCP tool catalog. /ledger#tool-surface - Worst database handler: 10.2 s -> 487 ms [reported]. Original resume reports organization-scoped, row-level-security-enforced sessions across a nine-organization Postgres estate, followed by query-plan fixes. The reported handler timings have no specified percentile or load. Source: Original resume: Hemut database hardening. /ledger#handler-latency - Parakeet digit WER after scorer correction: 0.303 WER -> 0.052 WER [measured]. Arithmetic means of the stored pk_wer and pk_wer2 scores for the 60 digit rows: 0.302882 and 0.052357. The same transcripts span six audio conditions. This measures a scoring correction, not new recognition performance; the number-aware scorer is not included in the retained probe. Source: parakeet_bench_results.json: digit-row scores. /ledger#digit-scoring - Paired probe WER across all conditions: 0.0486 Deepgram / 0.0537 Parakeet [measured]. Arithmetic means of stored dg_wer2 and pk_wer2 scores across all 240 paired rows. The pool includes 16 kHz clean audio, 8 kHz telephony, and four noise levels; it is not the separate 8 kHz summary or a production-call evaluation. Source: parakeet_bench_results.json: all-condition scores. /ledger#probe-wer - Name-clip WER with directory keyterms: 0.121 WER -> 0.014 WER [reported]. Historical report of the same seven synthesized name clips processed through 8 kHz mu-law audio, with and without directory keyterms. This is a name-subset comparison, not full-corpus WER or measured live-call routing success. Source: open-weight-stt-findings.md: name-clip comparison. /ledger#name-keyterms - Projected full-corpus WER with keyterms: 0.050 WER -> 0.033 WER [projected]. The historical report projects the seven-clip name improvement across the full corpus. The accompanying 204 ms latency is inherited from the Deepgram baseline, not a new timing measurement. Source: open-weight-stt-findings.md: summary table and projection footnote. /ledger#keyterms-projection - Article verdict categories: 4 categories [reported]. The original resume describes four verdict categories, a six-strategy fuzzy matcher, and four-level JSON repair. These are reported implementation features, not measured fact-checking accuracy. Source: Original resume: Veracity. /ledger#veracity-verdicts - Cache-hit energy: 98.6% modeled reduction: 0.007 kWh/query -> 0.0001 kWh/query [modeled]. Assumed energy constants: (1 - 0.0001 / 0.007) * 100 = 98.6%, rounded. The 0.386 kg CO2/kWh grid factor converts energy to carbon; it does not measure or validate energy consumption. No hardware energy measurement was made. Source: EcoPrompt energy model and original resume. /ledger#energy-cache - Right-sized energy: 90% modeled reduction: 0.007 kWh/query -> 0.0007 kWh/query [modeled]. Assumed per-query energy for the smaller versus larger model: (1 - 0.0007 / 0.007) * 100 = 90%. This is the model's right-sizing scenario, not measured energy or an observed average saving. Source: EcoPrompt model-routing energy assumptions and original resume. /ledger#energy-routing ## Contact and structured content Email: dongyeop0810@gmail.com GitHub: https://github.com/DY0810 LinkedIn: https://www.linkedin.com/in/dongyeopl Resume: /DongYeopLee_Resume.pdf JSON: /api/resume.json