Some results have saved runs. Some are reported from my work. Some are estimates. Here is the basis for each, including what the number does not tell you.
16 entries
Evidence entries
reported
Voice time to first speech
converse
1300 ms / to595 ms
Reported time-to-first-speech comparison following request parallelization, pooled HTTP connections, and end-of-turn tuning. The sampling window and percentile are unspecified.
Source
DongYeop Lee: Hemut voice latency
reported
Voice agent cost
converse
$0.10/min / to$0.037/min
Original resume reports a 63% reduction after moving STT, LLM, and TTS onto LiveKit's native Inference path and improving prompt caching. No matched billing window is attached to this comparison.
Source
Original resume: Hemut voice cost
reported
Prompt-cache hit rate
converse
95%
Hit rate reported in the original resume's voice-cost result. Its denominator and measurement window are unspecified; it is not interchangeable with a separate invoice's cached-token share.
Source
Original resume: Hemut prompt caching
reported
Document synchronization
converse
12 min / to1.05 s
Original resume compares the earlier Drive walk with the document-sync result. These are reported timings without an attached workload, sample size, or matched run.
Source
Original resume: Hemut document sync
reported
Document-sync cost reduction
converse
94%
Cost reduction reported alongside the document-sync timing in the original resume. The underlying billing window and workload are not specified.
Source
Original resume: Hemut document-sync cost
reported
Agent end-to-end accuracy
hemut-mcp
33% / to99%
Original resume outcome for the MCP agent layer and embedding-based tool retrieval. The reviewed baseline is recorded separately; no paired after-evaluation artifact is supplied for the 99% figure.
Source
Original resume: Hemut copilot accuracy
measured
Reviewed prompt baseline
hemut-mcp
32.8%
75 passing prompts divided by 229 reviewed prompts, rounded to one decimal place. This supports the baseline only, not a subsequent accuracy result.
Source
Internal prompt evaluation: reviewed June 16 run
Sample
75 passing / 229 reviewed prompts
Recorded
2026-06-16
reported
MCP tool catalog
hemut-mcp
192 tools
Tool catalog size reported in the original resume, with embedding-based retrieval selecting relevant tools. This is an implementation-scale figure, not an accuracy measurement or a current registry audit.
Source
Original resume: Hemut MCP tool catalog
reported
Worst database handler
hemut-mcp
10.2 s / to487 ms
Original resume reports organization-scoped, row-level-security-enforced sessions across a nine-organization Postgres estate, followed by query-plan fixes. The reported handler timings have no specified percentile or load.
Source
Original resume: Hemut database hardening
measured
Parakeet digit WER after scorer correction
benchmark
0.303 WER / to0.052 WER
Arithmetic means of the stored pk_wer and pk_wer2 scores for the 60 digit rows: 0.302882 and 0.052357. The same transcripts span six audio conditions. This measures a scoring correction, not new recognition performance; the number-aware scorer is not included in the retained probe.
Source
parakeet_bench_results.json: digit-row scores
Sample
10 synthesized digit clips across 6 conditions; 60 paired rows
Recorded
2026-08-05
measured
Paired probe WER across all conditions
benchmark
0.0486 Deepgram / 0.0537 Parakeet
Arithmetic means of stored dg_wer2 and pk_wer2 scores across all 240 paired rows. The pool includes 16 kHz clean audio, 8 kHz telephony, and four noise levels; it is not the separate 8 kHz summary or a production-call evaluation.
Source
parakeet_bench_results.json: all-condition scores
Sample
40 synthesized clips across 6 conditions; 240 paired rows, each with both model outputs
Recorded
2026-08-05
reported
Name-clip WER with directory keyterms
benchmark
0.121 WER / to0.014 WER
Historical report of the same seven synthesized name clips processed through 8 kHz mu-law audio, with and without directory keyterms. This is a name-subset comparison, not full-corpus WER or measured live-call routing success.
Source
open-weight-stt-findings.md: name-clip comparison
Sample
7 synthesized name clips
Recorded
2026-08-05
projected
Projected full-corpus WER with keyterms
benchmark
0.050 WER / to0.033 WER
The historical report projects the seven-clip name improvement across the full corpus. The accompanying 204 ms latency is inherited from the Deepgram baseline, not a new timing measurement.
Source
open-weight-stt-findings.md: summary table and projection footnote
Sample
40-clip corpus projection from a 7-clip name subset
Recorded
2026-08-05
reported
Article verdict categories
veracity
4 categories
The original resume describes four verdict categories, a six-strategy fuzzy matcher, and four-level JSON repair. These are reported implementation features, not measured fact-checking accuracy.
Source
Original resume: Veracity
modeled
Cache-hit energy: 98.6% modeled reduction
ecoprompt
0.007 kWh/query / to0.0001 kWh/query
Assumed energy constants: (1 - 0.0001 / 0.007) * 100 = 98.6%, rounded. The 0.386 kg CO2/kWh grid factor converts energy to carbon; it does not measure or validate energy consumption. No hardware energy measurement was made.
Source
EcoPrompt energy model and original resume
Sample
Cache-hit scenario, not an average across queries
modeled
Right-sized energy: 90% modeled reduction
ecoprompt
0.007 kWh/query / to0.0007 kWh/query
Assumed per-query energy for the smaller versus larger model: (1 - 0.0007 / 0.007) * 100 = 90%. This is the model's right-sizing scenario, not measured energy or an observed average saving.
Source
EcoPrompt model-routing energy assumptions and original resume