Can domain context improve local speech recognition?
AI.002 · Completed
Two quantized Whisper models made the same three command errors. Adding domain vocabulary to the smaller model corrected two of them on the same ten recordings.
Can domain context improve local speech recognition for Pixel’s Spanish voice commands on a Radeon RX 590?
01
Why this experiment
I wanted local speech recognition on the Fedora PC using hardware I already owned: an XFX Radeon RX 590 FATBOY with 8 GB of VRAM. Pixel’s commands include names such as Meteolab and Docker, so preserving those terms matters more than producing plausible Spanish.
When the smaller model missed those terms, I compared a larger model and then tested an initial prompt containing the vocabulary of the lab.
02
Success criteria
Assess whether each of ten recorded commands retains its intended meaning and critical entities. Compare reported average and median STT latency, loading time, recorded GPU memory and crashes.
This is a retrospective account. No numerical acceptance threshold or exact rule for reducing three repetitions to one utterance classification is preserved.
03
Environment and constraints
One speaker, ten Spanish recordings, three runs per recording and condition. The prompt was chosen after observing baseline failures and tested on the same audio, without a held-out set. Small and medium also use different quantizations.
Original conversation tables survive; raw logs, individual timings, complete transcripts, exact timing and memory instruments, and the repetition-to-PASS/FAIL rule are unavailable. WER, CER and end-to-end Pixel Voice latency were not measured.
04
Technical target
Run whisper.cpp locally through Vulkan on the RX 590, compare small-q5_1 with medium-q5_0, then test small-q5_1 with a fixed domain prompt on the same ten WAV files.
Hypothesis
A small Whisper model with domain vocabulary could preserve the important terms in these commands while retaining practical local transcription speed. A larger model was the alternative tested first.
Development
Lab notebook
13 Sep 2026Local speech on existing hardware
Pixel needed to recognize the Spanish commands I actually use in Afterlab. The Fedora PC already had an XFX Radeon RX 590 FATBOY, an older Polaris GPU with 8 GB of VRAM. The practical question was whether it could support useful local transcription, and whether errors called for a larger model or better domain context.
This page reconstructs a completed test from summary tables preserved in the original conversation. The raw run files have not been recovered. Figures reproduce those summaries and selected error fragments; they do not reconstruct missing observations.
Setup and corpus
The recorded environment was Fedora 44, the amdgpu kernel driver, Mesa 26.1.8 and Vulkan API 1.4.354. whisper.cpp 1.9.4-dev, commit 1da4dc8, was built with Vulkan and identified the GPU as AMD Radeon RX 590 Series (RADV POLARIS10). This test used Vulkan; neither OpenCL nor ROCm was the inference backend.
The CLI compared ggml-small-q5_1.bin and ggml-medium-q5_0.bin with Spanish selected, four threads and no timestamps: -l es -t 4 -nt. This was not a permanently loaded whisper-server. Loading was reported separately from STT timing.
The corpus contained ten recordings of my own voice, as mono 16 kHz, 16-bit PCM WAV files. An initial recording timeout problem produced some clips around twenty seconds long; it was corrected before the authoritative run. Final clip durations are not preserved, so no real-time factor is reported.
ID
Expected command
01
Pixel, pon el modo focus
02
Pixel, vuelve al modo normal
03
¿Cómo está el laboratorio?
04
¿Cómo está Docker ahora mismo?
05
¿Qué dispositivos están conectados?
06
Hola Pixel, ¿qué tal estás?
07
¿Qué temperatura tiene Meteolab?
08
Dime si hay alguna web caída.
09
Quiero saber cómo está funcionando el ordenador del laboratorio.
10
Pixel, revisa el estado del sistema y dime si ves algún problema.
Each condition used ten audios and three runs per audio: thirty inferences for small, thirty for medium, and thirty for small with a domain prompt. The prompt rerun reused exactly the same WAV files. The three conditions therefore account for ninety executions, but only ten unique utterances.
First comparison — small versus medium
Baseline summaries retained from the conversation. Thirty inferences per model, ten unique utterances. Memory measurement method and exact STT timer are not preserved.
Recorded metric
small-q5_1
medium-q5_0
Model loading
180 ms
297 ms
Average STT
810 ms
2,172 ms
Median / P50 STT
784 ms
2,169 ms
Recorded GPU memory
~189 MB
~539 MB
Functional classifications
7/10
7/10
Failed utterances
04, 07, 08
04, 07, 08
Crashes
0
0
Medium took approximately 2.7 times the average STT time and had approximately 2.9 times the recorded GPU memory, without improving the functional classifications. This compares these two quantized configurations on this corpus; it does not isolate model size from quantization or establish a general ranking of Whisper models.
Second comparison — domain context for small
After inspecting the baseline failures, I kept small-q5_1 and supplied an initial prompt with lab vocabulary:
Pixel, Afterlab, Meteolab, Docker, WordPress, CrowPanel, Fedora, Lab PC, modo focus, modo normal, webs, dispositivos.
The recorded configuration used --prompt and --carry-initial-prompt, with carry_initial_prompt=true. Medium was not rerun with the prompt.
Same ten WAV files, three runs each. Prompt selected after baseline errors; the exact repetition-to-classification rule is not preserved. Lower latency is an observation, not an established effect of prompting.
Recorded metric
Small baseline
Small + domain prompt
Average STT
810 ms
638 ms
Median / P50 STT
784 ms
610 ms
Functional classifications
7/10
9/10
Failed utterances
04, 07, 08
04
Crashes
0
0
GPU memory
~189 MB
Not re-reported
The prompted run had lower recorded latency. The experiment was designed to investigate recognition errors, not to isolate the cause of a performance difference, so I do not attribute that reduction to the prompt. No latency penalty was apparent in the retained summaries.
What changed in the transcripts
Selected retained fragments and utterance-level classifications. The missing baseline transcript for case 08 is not reconstructed.
The retained fragments show small transcribing Docker as doque both before and after prompting. Meteolab changed from el lab to Meteolab. The web caída command also changed from FAIL to PASS, with web caída retained in the prompted output. The full baseline transcript for that command is not preserved here.
These are fragments from recorded error analysis, not complete transcripts. Cases 07 and 08 account for the reported change from 7/10 to 9/10. Docker remains a known unresolved entity.
How to read the results
Functional classifications, not WER: each of the ten utterances received a PASS or FAIL based on preserving the intent and critical terms. Exact punctuation or string matching was not required. The rule used to combine the three repetitions into one classification is not preserved. Consequently, 7/10 cannot be converted into 21/30 successful transcriptions, and 9/10 cannot be converted into 27/30.
Recorded timings and memory: average and P50 come from the conversation’s original tables, which describe STT separately from loading. The timing implementation and memory measurement command are unavailable. The GPU figures are approximately 189 MB and 539 MB as reported; they are not a verified peak-memory comparison. Individual observations cannot currently be used to recalculate the aggregates.
Same-audio follow-up: domain terms were selected after observing failures and evaluated on the same ten recordings. The result shows recovery of two known errors on those recordings, not an independent estimate of generalization. There was no medium-with-prompt condition, CPU comparison, wake-word test or Fedora-versus-CrowPanel benchmark.
Decision and evidence retained
The selected configuration was small-q5_1 with Vulkan and the domain prompt. It retained the smaller model’s practical speed while improving the recorded classifications on the test commands. This documents the configuration decision; it does not establish full Pixel Voice deployment or end-to-end performance.
The authoritative baseline run was referenced as ~/Projects/whisper.cpp/afterlab-benchmark/results/run-20260913T083214Z/. The corpus was referenced under ~/Projects/whisper.cpp/afterlab-benchmark/, with files 01.wav through 10.wav. These are historical paths, not downloadable evidence links; the raw files have not been recovered and no separate prompt-run directory is preserved.
What survives is the benchmark protocol, the two summary tables, the domain prompt, final utterance classifications and selected transcript fragments. The next useful test is a new audio set with archived per-run outputs and explicit scoring, rather than another pass over the same known failures.
Findings
What the evidence says
01
Results
Both baseline models were recorded as functionally correct on 7/10 utterances. Small reported 810 ms average STT and approximately 189 MB GPU memory; medium reported 2,172 ms and approximately 539 MB. Neither preserved the critical terms in cases 04, 07 and 08.
Small with domain context was classified correct on 9/10 utterances, recovering Meteolab and web caída. Docker remained unresolved. Its reported average STT was 638 ms, with zero crashes; memory was not re-reported for this condition.
02
Hypothesis evaluation
Partially supported
03
What I learned
For these recordings, increasing model size did not improve the reported functional classifications. Adding vocabulary to the small model corrected two observed failures. That supports the local configuration choice, but does not establish accuracy on unseen speech.
Keep the benchmark’s evidence trail: per-run transcripts, timing fields, memory measurement commands and an explicit scoring rule. A summary can preserve a useful engineering decision while leaving important uncertainty.
04
Next steps
Evaluate the selected configuration on new recordings before making a general claim. Fix an explicit scoring rule, retain per-run outputs, and add word-level metrics alongside functional judgments.
Measure a persistent runtime separately from the CLI benchmark, including end-to-end latency, silence and unrelated speech, different speakers and ambient noise. Test medium with the same prompt if comparing the effect of model size under domain context.