A planner can return perfect JSON and still choose the wrong evidence. Two Afterlab benchmarks test where model-led source selection works—and where it needs an explicit boundary.
Can an AI planner reliably choose the evidence sources needed to answer operational questions about a personal lab?
01
Why this experiment
Pixel needs to answer questions that span current system state, recorded events and behaviour over time. Choosing the wrong source can produce a confident answer with missing evidence, even when every tool ID and JSON field is valid.
I wanted to measure that selection step independently, preserve failures and use the results to decide what the production planner should be allowed to handle.
02
Success criteria
For the final V1.1 comparison and extended test: strong GO requires at least 90% correct decisions, at least 95% valid JSON and p95 no higher than 2.5 s. A pilot requires at least 80% correct decisions and p95 no higher than 3.5 s, with the same JSON threshold.
The extended test additionally stops on a repeated critical confusion. Expected selections, acceptable alternatives and criteria are fixed before measuring. JSON validity, semantic correctness and latency are scored separately.
03
Environment and constraints
Requests ran through Afterlab’s miniPC and Open WebUI endpoint using the configured model IDs. This is a benchmark of that deployment path, not proof that every model ran locally or a general ranking of model quality.
The baseline uses 25 unique questions; Gemma has two passes. The extended source-selection test uses 30 different questions, a changed admission instruction and two diagnostic catalogue additions. Diagnostic timeout: 10 s. No tool execution or final-answer evaluation in the extended benchmark.
04
Technical target
Select zero to three permitted source IDs with one planner call. Reject malformed output, unknown IDs, duplicate sources and over-budget plans. Preserve direct routes for simple queries, explicit source labels and fallback.
Test whether current state, event history and sampled operational history remain distinct when the catalogue grows.
Hypothesis
A constrained planner with a small catalogue, a strict output contract and a maximum of three read-only sources can select the right operational evidence often enough to support a bounded pilot.
Development
Lab notebook
03 Oct 2026Establishing the planning baseline
Pixel already had direct routes for simple questions. The experiment concerned the harder cases: selecting a small set of evidence sources for a compound operational question. The planner returned tool IDs, not an answer. It could select at most three read-only sources or defer to existing routing.
The V1.1 benchmark used a frozen corpus of 25 unique questions, including negative cases where the planner should defer. Every model used the same prompt, catalogue, examples and strict parser. Correct decisions include an expected defer; an accepted JSON plan is not automatically a correct decision.
V1.1: same 25-question corpus. Gemma aggregate contains two passes. Latency is p95 of successful completions, not end-to-end answer latency.
Configured model / run
Benchmark result
Gemma 31B · two runs
46/50 · 92% correct JSON 50/50 p95 1.035 s GO · aggregated benchmark
GPT OSS 20B · prior reference
21/25 · 84% correct JSON 25/25 p95 4.881 s STOP
GPT OSS 120B · prior reference
19/25 · 76% correct JSON 25/25 p95 1.852 s STOP
Nemotron Nano · one run
19/25 · 76% correct JSON 25/25 p95 4.226 s STOP
Nemotron Super · one run
17/25 · 68% correct JSON 25/25 p95 3.774 s STOP
Same V1.1 corpus. GPT OSS rows are earlier reference runs. Model labels are the IDs configured in Afterlab; this is not a general model leaderboard.
Gemma was repeated because its first run combined 23/25 correct decisions with two latency spikes. Both runs were retained: the first had a p95 of 4.053 s, the second 0.820 s. The aggregate contains 46/50 correct decisions, a p95 of 1.035 s and a maximum of 7.841 s. The two runs repeat the same 25 questions; they are not 50 unique cases.
The proposed 1.6 s production budget changes the practical result. A replay of the recorded durations delivers 44/50 correct decisions within that budget, with two calls falling back. This is a projection, not a new live cancellation test. An aggregate GO at a 10 s diagnostic timeout does not guarantee 90% correct decisions delivered under the shorter budget.
DeepSeek requests were rejected by the provider, so no generation-quality score is assigned. Nemotron Ultra returned three correct completions, then timeouts and HTTP errors prevented a complete evaluation. It is excluded from the completed-model chart; 3/25 is not its semantic accuracy.
04 Oct 2026Testing the boundaries of operational history
The next test added operational_history and system_status to an isolated diagnostic catalogue. Production routing stayed intact. This new corpus contained 30 unique questions and explicitly measured source selection, including single-source questions. It did not execute tools or measure final-answer quality.
These are different tasks and corpora. The 92% baseline and the 63.33% extended result are not a before/after accuracy comparison on identical questions. The extended test asks whether the planner can respect additional semantic boundaries.
Operational History diagnostic: a different 30-question source-selection corpus. Counts come directly from the frozen per-case evidence.
Gemma returned valid JSON for all 30 cases, with no HTTP errors or diagnostic timeouts. Its p95 was 659.768 ms. Only 19/30 source selections were correct. Four cases confused Activity with Operational History; other failures substituted a live snapshot for historical behaviour or added irrelevant evidence.
Where the sources went wrong
These are actual failed selections from the frozen corpus. Questions retain their original Spanish wording; tool IDs retain their exact names.
For MQTT stability, an event log cannot substitute for recorded runtime transitions. For a question about events and miniPC behaviour, today’s system snapshot cannot substitute for telemetry over that period. A syntactically valid answer can therefore omit precisely the evidence the question requires.
The extended gate required at least 90% correct decisions for a strong GO or 80% for a pilot, with JSON and latency thresholds. A repeated critical confusion also triggered STOP. This run failed on semantic accuracy and on the repeated critical pattern. The prompt and expected answers were not adjusted after observing the errors; no favourable second pass was selected.
04 Oct 2026Keeping the operational boundary explicit
The decision was to keep Operational History outside the production Composite Planner. A closed deterministic composition layer was implemented to combine current state, recorded events and historical telemetry only when the question needs them. It reuses the period parser, caps actual reads at three and feeds the existing synthesis path.
Implemented alternative: explicit source boundaries, at most three reads, one existing synthesis path. Live checkpoint pending.
Simple deterministic reads take precedence, followed by operational composition, the existing Composite Planner and the existing fallback. Historical observations enter HISTORY; live readings remain current facts. Composition bypasses planner counters and does not add another LLM planning call.
Validation boundary: the composition implementation passed its local routing and regression tests. Its deployment and live behavioural checkpoint are still pending in the current project documentation. These tests do not establish superior live accuracy or latency against the planner benchmark. Historical telemetry itself is separately documented as running in production.
Evidence and reproducibility
Both benchmark protocols preserve prompt and corpus fingerprints, expected selections, per-call timings and decision criteria. Latency uses successful completions and nearest-rank p95. The extended evidence contains 30 case records; the figures here are derived from those records, with baseline model summaries taken from the frozen report.
The V1.1 Gemma aggregate reached 46/50 correct decisions and 50/50 valid JSON, with p95 1.035 s. The extended test reached 19/30 correct decisions despite 30/30 valid JSON and p95 659.768 ms, so the proposed catalogue expansion received STOP.
Operational History remains outside the production Composite Planner. A deterministic alternative is implemented and tested; its live checkpoint remains pending.
02
Hypothesis evaluation
Partially supported
03
What I learned
Valid structure does not establish correct evidence selection. Good latency does not rescue repeated semantic errors. A successful small-catalogue benchmark does not establish that additional, closely related sources can safely be added.
Keep current facts, recorded events and telemetry over a period explicit. Use model planning within an evaluated scope, and validate an expanded scope before changing production.
04
Next steps
Complete the deterministic composition checkpoint with real operational questions and inspect source labels, periods, partial-data behaviour and planner bypass. Record latency and selection outcomes without presenting local test passes as live benchmark accuracy.
Any future planner expansion needs a new versioned corpus and predeclared criteria. Preserve the original STOP result.