Skip to content

Can domain context improve local speech recognition?

AI.002 · Completed

Two quantized Whisper models made the same three command errors. Adding domain vocabulary to the smaller model corrected two of them on the same ten recordings.

Status
Completed
Started
September 2026
Completed
September 2026
Topics
AIExperimental computing
Tools & Stack
Fedora Vulkan whisper.cpp

Research brief

The question

Can domain context improve local speech recognition for Pixel’s Spanish voice commands on a Radeon RX 590?

Why this experiment

I wanted local speech recognition on the Fedora PC using hardware I already owned: an XFX Radeon RX 590 FATBOY with 8 GB of VRAM. Pixel’s commands include names such as Meteolab and Docker, so preserving those terms matters more than producing plausible Spanish.

When the smaller model missed those terms, I compared a larger model and then tested an initial prompt containing the vocabulary of the lab.

Success criteria

Assess whether each of ten recorded commands retains its intended meaning and critical entities. Compare reported average and median STT latency, loading time, recorded GPU memory and crashes.

This is a retrospective account. No numerical acceptance threshold or exact rule for reducing three repetitions to one utterance classification is preserved.

Environment and constraints

One speaker, ten Spanish recordings, three runs per recording and condition. The prompt was chosen after observing baseline failures and tested on the same audio, without a held-out set. Small and medium also use different quantizations.

Original conversation tables survive; raw logs, individual timings, complete transcripts, exact timing and memory instruments, and the repetition-to-PASS/FAIL rule are unavailable. WER, CER and end-to-end Pixel Voice latency were not measured.

Technical target

Run whisper.cpp locally through Vulkan on the RX 590, compare small-q5_1 with medium-q5_0, then test small-q5_1 with a fixed domain prompt on the same ten WAV files.

Hypothesis

A small Whisper model with domain vocabulary could preserve the important terms in these commands while retaining practical local transcription speed. A larger model was the alternative tested first.

Development

Lab notebook

13 Sep 2026Local speech on existing hardware

Pixel needed to recognize the Spanish commands I actually use in Afterlab. The Fedora PC already had an XFX Radeon RX 590 FATBOY, an older Polaris GPU with 8 GB of VRAM. The practical question was whether it could support useful local transcription, and whether errors called for a larger model or better domain context.

This page reconstructs a completed test from summary tables preserved in the original conversation. The raw run files have not been recovered. Figures reproduce those summaries and selected error fragments; they do not reconstruct missing observations.

Setup and corpus

The recorded environment was Fedora 44, the amdgpu kernel driver, Mesa 26.1.8 and Vulkan API 1.4.354. whisper.cpp 1.9.4-dev, commit 1da4dc8, was built with Vulkan and identified the GPU as AMD Radeon RX 590 Series (RADV POLARIS10). This test used Vulkan; neither OpenCL nor ROCm was the inference backend.

The CLI compared ggml-small-q5_1.bin and ggml-medium-q5_0.bin with Spanish selected, four threads and no timestamps: -l es -t 4 -nt. This was not a permanently loaded whisper-server. Loading was reported separately from STT timing.

The corpus contained ten recordings of my own voice, as mono 16 kHz, 16-bit PCM WAV files. An initial recording timeout problem produced some clips around twenty seconds long; it was corrected before the authoritative run. Final clip durations are not preserved, so no real-time factor is reported.

IDExpected command
01Pixel, pon el modo focus
02Pixel, vuelve al modo normal
03¿Cómo está el laboratorio?
04¿Cómo está Docker ahora mismo?
05¿Qué dispositivos están conectados?
06Hola Pixel, ¿qué tal estás?
07¿Qué temperatura tiene Meteolab?
08Dime si hay alguna web caída.
09Quiero saber cómo está funcionando el ordenador del laboratorio.
10Pixel, revisa el estado del sistema y dime si ves algún problema.

Each condition used ten audios and three runs per audio: thirty inferences for small, thirty for medium, and thirty for small with a domain prompt. The prompt rerun reused exactly the same WAV files. The three conditions therefore account for ninety executions, but only ten unique utterances.

First comparison — small versus medium

Small-q5_1 and medium-q5_0 both classified correct on 7 of 10 utterances. Average STT 810 versus 2172 milliseconds; recorded GPU memory approximately 189 versus 539 megabytes.
Baseline summaries retained from the conversation. Thirty inferences per model, ten unique utterances. Memory measurement method and exact STT timer are not preserved.
Recorded metricsmall-q5_1medium-q5_0
Model loading180 ms297 ms
Average STT810 ms2,172 ms
Median / P50 STT784 ms2,169 ms
Recorded GPU memory~189 MB~539 MB
Functional classifications7/107/10
Failed utterances04, 07, 0804, 07, 08
Crashes00

Medium took approximately 2.7 times the average STT time and had approximately 2.9 times the recorded GPU memory, without improving the functional classifications. This compares these two quantized configurations on this corpus; it does not isolate model size from quantization or establish a general ranking of Whisper models.

Second comparison — domain context for small

After inspecting the baseline failures, I kept small-q5_1 and supplied an initial prompt with lab vocabulary:

Pixel, Afterlab, Meteolab, Docker, WordPress, CrowPanel, Fedora, Lab PC, modo focus, modo normal, webs, dispositivos.

The recorded configuration used --prompt and --carry-initial-prompt, with carry_initial_prompt=true. Medium was not rerun with the prompt.

Small baseline: 7 of 10 utterances correct, 810 milliseconds average and 784 milliseconds median. Small with domain prompt: 9 of 10, 638 milliseconds average and 610 milliseconds median. Same recordings, no held-out set.
Same ten WAV files, three runs each. Prompt selected after baseline errors; the exact repetition-to-classification rule is not preserved. Lower latency is an observation, not an established effect of prompting.
Recorded metricSmall baselineSmall + domain prompt
Average STT810 ms638 ms
Median / P50 STT784 ms610 ms
Functional classifications7/109/10
Failed utterances04, 07, 0804
Crashes00
GPU memory~189 MBNot re-reported

The prompted run had lower recorded latency. The experiment was designed to investigate recognition errors, not to isolate the cause of a performance difference, so I do not attribute that reduction to the prompt. No latency penalty was apparent in the retained summaries.

What changed in the transcripts

Recorded error fragments: Docker remains doque and FAIL. Meteolab changes from el lab and FAIL to Meteolab and PASS. Web caída changes from FAIL to PASS. Complete transcripts are not available.
Selected retained fragments and utterance-level classifications. The missing baseline transcript for case 08 is not reconstructed.

The retained fragments show small transcribing Docker as doque both before and after prompting. Meteolab changed from el lab to Meteolab. The web caída command also changed from FAIL to PASS, with web caída retained in the prompted output. The full baseline transcript for that command is not preserved here.

These are fragments from recorded error analysis, not complete transcripts. Cases 07 and 08 account for the reported change from 7/10 to 9/10. Docker remains a known unresolved entity.

How to read the results

Functional classifications, not WER: each of the ten utterances received a PASS or FAIL based on preserving the intent and critical terms. Exact punctuation or string matching was not required. The rule used to combine the three repetitions into one classification is not preserved. Consequently, 7/10 cannot be converted into 21/30 successful transcriptions, and 9/10 cannot be converted into 27/30.

Recorded timings and memory: average and P50 come from the conversation’s original tables, which describe STT separately from loading. The timing implementation and memory measurement command are unavailable. The GPU figures are approximately 189 MB and 539 MB as reported; they are not a verified peak-memory comparison. Individual observations cannot currently be used to recalculate the aggregates.

Same-audio follow-up: domain terms were selected after observing failures and evaluated on the same ten recordings. The result shows recovery of two known errors on those recordings, not an independent estimate of generalization. There was no medium-with-prompt condition, CPU comparison, wake-word test or Fedora-versus-CrowPanel benchmark.

Decision and evidence retained

The selected configuration was small-q5_1 with Vulkan and the domain prompt. It retained the smaller model’s practical speed while improving the recorded classifications on the test commands. This documents the configuration decision; it does not establish full Pixel Voice deployment or end-to-end performance.

The authoritative baseline run was referenced as ~/Projects/whisper.cpp/afterlab-benchmark/results/run-20260913T083214Z/. The corpus was referenced under ~/Projects/whisper.cpp/afterlab-benchmark/, with files 01.wav through 10.wav. These are historical paths, not downloadable evidence links; the raw files have not been recovered and no separate prompt-run directory is preserved.

What survives is the benchmark protocol, the two summary tables, the domain prompt, final utterance classifications and selected transcript fragments. The next useful test is a new audio set with archived per-run outputs and explicit scoring, rather than another pass over the same known failures.

Findings

What the evidence says

Results

Both baseline models were recorded as functionally correct on 7/10 utterances. Small reported 810 ms average STT and approximately 189 MB GPU memory; medium reported 2,172 ms and approximately 539 MB. Neither preserved the critical terms in cases 04, 07 and 08.

Small with domain context was classified correct on 9/10 utterances, recovering Meteolab and web caída. Docker remained unresolved. Its reported average STT was 638 ms, with zero crashes; memory was not re-reported for this condition.

Hypothesis evaluation

Partially supported

What I learned

For these recordings, increasing model size did not improve the reported functional classifications. Adding vocabulary to the small model corrected two observed failures. That supports the local configuration choice, but does not establish accuracy on unseen speech.

Keep the benchmark’s evidence trail: per-run transcripts, timing fields, memory measurement commands and an explicit scoring rule. A summary can preserve a useful engineering decision while leaving important uncertainty.

Next steps

Evaluate the selected configuration on new recordings before making a general claim. Fix an explicit scoring rule, retain per-run outputs, and add word-level metrics alongside functional judgments.

Measure a persistent runtime separately from the CLI benchmark, including end-to-end latency, silence and unrelated speech, different speakers and ambient noise. Test medium with the same prompt if comparing the effect of model size under domain context.

Other experiments

All experiments
  1. Can an AI planner choose the right evidence?

    AI.003
    Agents · AI · Experimental computing
    Completed October 2026
  2. Training a language model from scratch on 4GB of VRAM

    AI.001
    AI · Language Models · Local AI
    Completed October 2026

Part of a project · Personal lab / AI & automation platform

Afterlab

This experiment is one piece of a bigger build. See how it fits into the project, what it fed into and where the idea went next.