Skip to content

Can an AI planner choose the right evidence?

AI.003 · Completed

A planner can return perfect JSON and still choose the wrong evidence. Two Afterlab benchmarks test where model-led source selection works—and where it needs an explicit boundary.

Status
Completed
Started
October 2026
Completed
October 2026
Topics
AgentsAIExperimental computing
Tools & Stack
Docker Ollama Open WebUI Python

Research brief

The question

Can an AI planner reliably choose the evidence sources needed to answer operational questions about a personal lab?

Why this experiment

Pixel needs to answer questions that span current system state, recorded events and behaviour over time. Choosing the wrong source can produce a confident answer with missing evidence, even when every tool ID and JSON field is valid.

I wanted to measure that selection step independently, preserve failures and use the results to decide what the production planner should be allowed to handle.

Success criteria

For the final V1.1 comparison and extended test: strong GO requires at least 90% correct decisions, at least 95% valid JSON and p95 no higher than 2.5 s. A pilot requires at least 80% correct decisions and p95 no higher than 3.5 s, with the same JSON threshold.

The extended test additionally stops on a repeated critical confusion. Expected selections, acceptable alternatives and criteria are fixed before measuring. JSON validity, semantic correctness and latency are scored separately.

Environment and constraints

Requests ran through Afterlab’s miniPC and Open WebUI endpoint using the configured model IDs. This is a benchmark of that deployment path, not proof that every model ran locally or a general ranking of model quality.

The baseline uses 25 unique questions; Gemma has two passes. The extended source-selection test uses 30 different questions, a changed admission instruction and two diagnostic catalogue additions. Diagnostic timeout: 10 s. No tool execution or final-answer evaluation in the extended benchmark.

Technical target

Select zero to three permitted source IDs with one planner call. Reject malformed output, unknown IDs, duplicate sources and over-budget plans. Preserve direct routes for simple queries, explicit source labels and fallback.

Test whether current state, event history and sampled operational history remain distinct when the catalogue grows.

Hypothesis

A constrained planner with a small catalogue, a strict output contract and a maximum of three read-only sources can select the right operational evidence often enough to support a bounded pilot.

03 Oct 2026Establishing the planning baseline

Pixel already had direct routes for simple questions. The experiment concerned the harder cases: selecting a small set of evidence sources for a compound operational question. The planner returned tool IDs, not an answer. It could select at most three read-only sources or defer to existing routing.

The V1.1 benchmark used a frozen corpus of 25 unique questions, including negative cases where the planner should defer. Every model used the same prompt, catalogue, examples and strict parser. Correct decisions include an expected defer; an accepted JSON plan is not automatically a correct decision.

V1.1 source-selection benchmark: Gemma 46 of 50 correct; GPT OSS 20B 21 of 25; GPT OSS 120B and Nemotron Nano 19 of 25; Nemotron Super 17 of 25.
V1.1: same 25-question corpus. Gemma aggregate contains two passes. Latency is p95 of successful completions, not end-to-end answer latency.
Configured model / runBenchmark result
Gemma 31B · two runs46/50 · 92% correct
JSON 50/50
p95 1.035 s
GO · aggregated benchmark
GPT OSS 20B · prior reference21/25 · 84% correct
JSON 25/25
p95 4.881 s
STOP
GPT OSS 120B · prior reference19/25 · 76% correct
JSON 25/25
p95 1.852 s
STOP
Nemotron Nano · one run19/25 · 76% correct
JSON 25/25
p95 4.226 s
STOP
Nemotron Super · one run17/25 · 68% correct
JSON 25/25
p95 3.774 s
STOP
Same V1.1 corpus. GPT OSS rows are earlier reference runs. Model labels are the IDs configured in Afterlab; this is not a general model leaderboard.

Gemma was repeated because its first run combined 23/25 correct decisions with two latency spikes. Both runs were retained: the first had a p95 of 4.053 s, the second 0.820 s. The aggregate contains 46/50 correct decisions, a p95 of 1.035 s and a maximum of 7.841 s. The two runs repeat the same 25 questions; they are not 50 unique cases.

The proposed 1.6 s production budget changes the practical result. A replay of the recorded durations delivers 44/50 correct decisions within that budget, with two calls falling back. This is a projection, not a new live cancellation test. An aggregate GO at a 10 s diagnostic timeout does not guarantee 90% correct decisions delivered under the shorter budget.

DeepSeek requests were rejected by the provider, so no generation-quality score is assigned. Nemotron Ultra returned three correct completions, then timeouts and HTTP errors prevented a complete evaluation. It is excluded from the completed-model chart; 3/25 is not its semantic accuracy.

04 Oct 2026Testing the boundaries of operational history

The next test added operational_history and system_status to an isolated diagnostic catalogue. Production routing stayed intact. This new corpus contained 30 unique questions and explicitly measured source selection, including single-source questions. It did not execute tools or measure final-answer quality.

These are different tasks and corpora. The 92% baseline and the 63.33% extended result are not a before/after accuracy comparison on identical questions. The extended test asks whether the planner can respect additional semantic boundaries.

Extended benchmark: 19 of 30 correct selections, 30 of 30 valid JSON, p95 659.768 milliseconds. Correct selections by class: A 6/10, B 4/5, C 3/4, D 3/5, E 1/2, F 2/2, G 0/2.
Operational History diagnostic: a different 30-question source-selection corpus. Counts come directly from the frozen per-case evidence.

Gemma returned valid JSON for all 30 cases, with no HTTP errors or diagnostic timeouts. Its p95 was 659.768 ms. Only 19/30 source selections were correct. Four cases confused Activity with Operational History; other failures substituted a live snapshot for historical behaviour or added irrelevant evidence.

Where the sources went wrong

These are actual failed selections from the frozen corpus. Questions retain their original Spanish wording; tool IDs retain their exact names.

A06 — MQTT degradation

¿MQTT ha tenido degradaciones hoy?

Expected: operational_history
Selected: mission_control + activity_history

C02 — Current MQTT state and weekly stability

¿MQTT está bien ahora y ha sido estable esta semana?

Expected: mission_control + operational_history
Selected: mission_control + activity_history

D04 — Events and miniPC behaviour

¿Qué ha pasado hoy y cómo se ha comportado el miniPC?

Expected: activity_history + operational_history
Selected: activity_history + system_status

E01 — Events, current state and historical peaks

¿Qué ha pasado hoy, cómo está todo ahora y ha tenido el miniPC algún pico raro?

Expected: mission_control + activity_history + operational_history
Selected: mission_control + activity_history + system_status

For MQTT stability, an event log cannot substitute for recorded runtime transitions. For a question about events and miniPC behaviour, today’s system snapshot cannot substitute for telemetry over that period. A syntactically valid answer can therefore omit precisely the evidence the question requires.

The extended gate required at least 90% correct decisions for a strong GO or 80% for a pilot, with JSON and latency thresholds. A repeated critical confusion also triggered STOP. This run failed on semantic accuracy and on the repeated critical pattern. The prompt and expected answers were not adjusted after observing the errors; no favourable second pass was selected.

04 Oct 2026Keeping the operational boundary explicit

The decision was to keep Operational History outside the production Composite Planner. A closed deterministic composition layer was implemented to combine current state, recorded events and historical telemetry only when the question needs them. It reuses the period parser, caps actual reads at three and feeds the existing synthesis path.

Three evidence layers: LIVE for current state, Activity for recorded events, Operational History for behaviour over a period. Deterministic composition combines required sources, then uses existing synthesis.
Implemented alternative: explicit source boundaries, at most three reads, one existing synthesis path. Live checkpoint pending.

Simple deterministic reads take precedence, followed by operational composition, the existing Composite Planner and the existing fallback. Historical observations enter HISTORY; live readings remain current facts. Composition bypasses planner counters and does not add another LLM planning call.

Validation boundary: the composition implementation passed its local routing and regression tests. Its deployment and live behavioural checkpoint are still pending in the current project documentation. These tests do not establish superior live accuracy or latency against the planner benchmark. Historical telemetry itself is separately documented as running in production.

Evidence and reproducibility

Both benchmark protocols preserve prompt and corpus fingerprints, expected selections, per-call timings and decision criteria. Latency uses successful completions and nearest-rank p95. The extended evidence contains 30 case records; the figures here are derived from those records, with baseline model summaries taken from the frozen report.

Read the V1.1 benchmark report, the Operational History STOP decision and the per-case evidence. The implementation status distinguishes tested code from completed live validation.

Findings

What the evidence says

Results

The V1.1 Gemma aggregate reached 46/50 correct decisions and 50/50 valid JSON, with p95 1.035 s. The extended test reached 19/30 correct decisions despite 30/30 valid JSON and p95 659.768 ms, so the proposed catalogue expansion received STOP.

Operational History remains outside the production Composite Planner. A deterministic alternative is implemented and tested; its live checkpoint remains pending.

Hypothesis evaluation

Partially supported

What I learned

Valid structure does not establish correct evidence selection. Good latency does not rescue repeated semantic errors. A successful small-catalogue benchmark does not establish that additional, closely related sources can safely be added.

Keep current facts, recorded events and telemetry over a period explicit. Use model planning within an evaluated scope, and validate an expanded scope before changing production.

Next steps

Complete the deterministic composition checkpoint with real operational questions and inspect source labels, periods, partial-data behaviour and planner bypass. Record latency and selection outcomes without presenting local test passes as live benchmark accuracy.

Any future planner expansion needs a new versioned corpus and predeclared criteria. Preserve the original STOP result.

Other experiments

All experiments
  1. Can domain context improve local speech recognition?

    AI.002
    AI · Experimental computing
    Completed September 2026
  2. Training a language model from scratch on 4GB of VRAM

    AI.001
    AI · Language Models · Local AI
    Completed October 2026

Part of a project · Personal lab / AI & automation platform

Afterlab

This experiment is one piece of a bigger build. See how it fits into the project, what it fed into and where the idea went next.