Back to the work

Behind the numbers.

Some results have saved runs. Some are reported from my work. Some are estimates. Here is the basis for each, including what the number does not tell you.

16 entries

Evidence entries

reported

Voice time to first speech

1300 msto595 ms

Reported time-to-first-speech comparison following request parallelization, pooled HTTP connections, and end-of-turn tuning. The sampling window and percentile are unspecified.

Source
DongYeop Lee: Hemut voice latency
reported

Voice agent cost

$0.10/minto$0.037/min

Original resume reports a 63% reduction after moving STT, LLM, and TTS onto LiveKit's native Inference path and improving prompt caching. No matched billing window is attached to this comparison.

Source
Original resume: Hemut voice cost
reported

Prompt-cache hit rate

95%

Hit rate reported in the original resume's voice-cost result. Its denominator and measurement window are unspecified; it is not interchangeable with a separate invoice's cached-token share.

Source
Original resume: Hemut prompt caching
reported

Document synchronization

12 minto1.05 s

Original resume compares the earlier Drive walk with the document-sync result. These are reported timings without an attached workload, sample size, or matched run.

Source
Original resume: Hemut document sync
reported

Document-sync cost reduction

94%

Cost reduction reported alongside the document-sync timing in the original resume. The underlying billing window and workload are not specified.

Source
Original resume: Hemut document-sync cost
reported

Agent end-to-end accuracy

33%to99%

Original resume outcome for the MCP agent layer and embedding-based tool retrieval. The reviewed baseline is recorded separately; no paired after-evaluation artifact is supplied for the 99% figure.

Source
Original resume: Hemut copilot accuracy
measured

Reviewed prompt baseline

32.8%

75 passing prompts divided by 229 reviewed prompts, rounded to one decimal place. This supports the baseline only, not a subsequent accuracy result.

Source
Internal prompt evaluation: reviewed June 16 run
Sample
75 passing / 229 reviewed prompts
Recorded
2026-06-16
reported

MCP tool catalog

192 tools

Tool catalog size reported in the original resume, with embedding-based retrieval selecting relevant tools. This is an implementation-scale figure, not an accuracy measurement or a current registry audit.

Source
Original resume: Hemut MCP tool catalog
reported

Worst database handler

10.2 sto487 ms

Original resume reports organization-scoped, row-level-security-enforced sessions across a nine-organization Postgres estate, followed by query-plan fixes. The reported handler timings have no specified percentile or load.

Source
Original resume: Hemut database hardening
measured

Parakeet digit WER after scorer correction

0.303 WERto0.052 WER

Arithmetic means of the stored pk_wer and pk_wer2 scores for the 60 digit rows: 0.302882 and 0.052357. The same transcripts span six audio conditions. This measures a scoring correction, not new recognition performance; the number-aware scorer is not included in the retained probe.

Source
parakeet_bench_results.json: digit-row scores
Sample
10 synthesized digit clips across 6 conditions; 60 paired rows
Recorded
2026-08-05
measured

Paired probe WER across all conditions

0.0486 Deepgram / 0.0537 Parakeet

Arithmetic means of stored dg_wer2 and pk_wer2 scores across all 240 paired rows. The pool includes 16 kHz clean audio, 8 kHz telephony, and four noise levels; it is not the separate 8 kHz summary or a production-call evaluation.

Source
parakeet_bench_results.json: all-condition scores
Sample
40 synthesized clips across 6 conditions; 240 paired rows, each with both model outputs
Recorded
2026-08-05
reported

Name-clip WER with directory keyterms

0.121 WERto0.014 WER

Historical report of the same seven synthesized name clips processed through 8 kHz mu-law audio, with and without directory keyterms. This is a name-subset comparison, not full-corpus WER or measured live-call routing success.

Source
open-weight-stt-findings.md: name-clip comparison
Sample
7 synthesized name clips
Recorded
2026-08-05
projected

Projected full-corpus WER with keyterms

0.050 WERto0.033 WER

The historical report projects the seven-clip name improvement across the full corpus. The accompanying 204 ms latency is inherited from the Deepgram baseline, not a new timing measurement.

Source
open-weight-stt-findings.md: summary table and projection footnote
Sample
40-clip corpus projection from a 7-clip name subset
Recorded
2026-08-05
reported

Article verdict categories

4 categories

The original resume describes four verdict categories, a six-strategy fuzzy matcher, and four-level JSON repair. These are reported implementation features, not measured fact-checking accuracy.

Source
Original resume: Veracity
modeled

Cache-hit energy: 98.6% modeled reduction

0.007 kWh/queryto0.0001 kWh/query

Assumed energy constants: (1 - 0.0001 / 0.007) * 100 = 98.6%, rounded. The 0.386 kg CO2/kWh grid factor converts energy to carbon; it does not measure or validate energy consumption. No hardware energy measurement was made.

Source
EcoPrompt energy model and original resume
Sample
Cache-hit scenario, not an average across queries
modeled

Right-sized energy: 90% modeled reduction

0.007 kWh/queryto0.0007 kWh/query

Assumed per-query energy for the smaller versus larger model: (1 - 0.0007 / 0.007) * 100 = 90%. This is the model's right-sizing scenario, not measured energy or an observed average saving.

Source
EcoPrompt model-routing energy assumptions and original resume
Sample
Right-sized-query scenario