Session Search — Investigation

Problem scoping, measurements, and design rationale — companion to PR #26419 · for as-shipped behaviour, see results & value-prop →

We profiled session_search's end-to-end cost against a 280-session database to investigate three user-facing symptoms: searches taking 30 s+, empty results for queries that should match, and late-session content missing from summaries. The investigation confirmed that aux-LLM summarisation dominates latency by 1,299×–6,293× over snippets, falsified the "fast inflates the calling-agent context" assumption (summary blobs are 3–4× larger), and surfaced the deeper finding that fast and summary share the same FTS5 retrieval set — so retrieval quality, not synthesis, is the highest-leverage lever. This page is the investigation record that informed the toolkit eventually shipped.

4
Queries profiled
280
Sessions in DB
6,293×
Max latency ratio
3.4×
Blob-size ratio
3
Hypotheses tested
1. Problem

session_search felt slow and unreliable

Three symptoms surfaced from real recall workflows: searches taking 30 s+, occasional empty results for queries the user knew should match, and substantive late-session content missing from summaries. Each got a hypothesis to test against measurement.

H1 — "summarisation is slow"

Searches taking 30+ seconds. Each call fans out up to 3 parallel auxiliary LLM calls; the slowest dominates wall time.

H2 — "100K window misses late content"

Truncation centres on the first FTS5 match; if that's in message 1 (e.g. session-title query), late-session work isn't fed to the summariser.

H3 — "exact title returns nothing"

Some queries return zero results despite a known matching session. Suspect: sqlite3.OperationalError swallowed on FTS5 query parser failure.

2. System as it was

Three-stage pipeline; cost concentrated in stage 3

The pre-change session_search() orchestrated DB retrieval, in-process prep, and parallel auxiliary summarisation. Every call paid for stage 3 — there was no way to opt out of the LLM fan-out. This section captures the starting point against which the investigation measured.

Ctrl/Cmd + wheel to zoom · Drag when zoomed

Loading...
3. Findings

Latency & size measured against real DB

Single run per query. Auxiliary model: whatever auxiliary.session_search resolves to in the live config. DB: 247 MB SQLite snapshot, 280 sessions, 18,157 messages. These numbers tell us how much each mode costs; the consumer replay in §6 tells us what each mode delivers.

Per-query tool latency & payload size — three modes

Measures the tool's intrinsic latency profile. Fast is a single tool call. Summary is a single tool call but triggers an auxiliary LLM fan-out — its wall time is its end-to-end cost. Guided is never used in isolation — it rides on a prior fast call plus an agent-inference turn that picks the anchors; only the raw DB read is shown here. For full interaction timing (tool walls + main-agent reasoning + final response composition), see results.html, where v2 measured this across 5 iterations × 18 scenarios on the as-shipped toolkit.

Queryfast (1 call)summary (1 call)guided (DB read)fast Bsummary Bguided B
[q1 topic keywords]11 ms65.85 s~2 ms6,56821,389131,599
session_search_investigation5 ms63.71 s~2 ms6,54522,54661,787
daily plan7 ms42.41 s~2 ms5,61815,72131,808
atlas commons rewording0.27 ms0.17 ms— (no anchors)138141—

Steering-improvements run: the wider limit and multi-anchor guided shown here were the candidates measured in this investigation; both were promoted to the shipped toolkit (multi-anchor in production, limit cap raised to 10). The wider limit roughly doubles fast/summary payload and roughly triples guided payload. DB: 247 MB SQLite snapshot, 280 sessions, 18,157 messages.

Payload trade-off, sharpened: multi-anchor guided's payload is ~3–6× larger than summary's recap (131 kB vs 21 kB on Q1) because the agent gets three full message windows of raw transcript. Q1's 131 kB is dominated by JSON tool-result bodies the agent doesn't need for steering — a leaner-payload tool-noise filter on guided was a candidate follow-up surfaced by this investigation.

Q4 returns zero FTS5 hits (the wider limit doesn't help when nothing matches), so guided has nothing to anchor on. The harness correctly skips it; the deeper retrieval bug is the deeper lever this PR doesn't pull — see the closing structural insight in §6.

H1 confirmed: summarisation is the dominant cost
  • 99%+ of summary-mode latency is in _summarize_session
  • DB / dedupe / format / truncate together: < 15 ms
  • LLM auxiliary call: 27–37 s wall, 53–72 s sum-of-parallel-walls
"Fast inflates the calling-agent context" — falsified
  • This was our argument against fast-as-default early in planning
  • Reality: summary blobs are 2.9×–3.9× larger than fast
  • The LLM, asked to be "thorough but concise", expands rather than compresses
H2 untested in this run
  • Affects summary mode only; fast doesn't fetch transcripts
  • If summary stays niche, H2 demotes to a follow-up
H3 not reproduced today; Q4 anomaly explained
  • Q2 was the title-shape test; succeeded in both modes
  • Q4 returned 0 hits in both — FTS5 default-AND tokenisation footgun
  • H3 likely needs specific FTS5-special chars to trigger
4. State walk

Q1 ([q1 topic keywords]) traced through both modes

The numbers tell you which mode is faster; the state walk tells you which mode answers the user's question. Showing Result 1 of 3 for length.

STEP 1 — INPUT (shared)
The agent calls session_search
SHARED
{
  "query": "[q1 topic keywords]",
  "limit": 3,
  "mode": "summary"  // or "fast"
}
STEP 2 — FTS5 RAW HITS (shared)
db.search_messages() — up to 50 ranked message hits
FTS5 BM25 against messages_fts (unicode61). Each row carries snippet, +/-1 message context, session metadata. ~10 ms wall.
SHARED — one row of ~50
{
  "id": 1234567,
  "session_id": "cron_[sid_A]_20260501_1600",
  "role": "assistant",
  "snippet": "...>>>[q1]<<<->>>[topic]<<<->>>[keywords]<<< Task Created\nThe user's **#1 priority for May 1, 2026** ...",
  "content": "(full message text — typically 1-10 KB)",
  "timestamp": 1748793600.123,
  "tool_name": null,
  "source": "cron",
  "model": "claude-opus-4-6",
  "session_started": 1748793603.456,
  "context": [
    {"role": "assistant", "content": "Now let me search for each session ID..."},
    {"role": "tool", "content": "{...truncated...}"}
  ]
}
STEP 3 — DEDUPED SESSIONS (shared)
Walk parent_session_id chain, dedupe, top 3 unique
SHARED
[
  {"session_id": "cron_[sid_A]_20260501_1600"},  // 16:00 cron check-in
  {"session_id": "cron_[sid_B]_20260501_2330"},  // 23:30 EOD wrap-up
  {"session_id": "[earlier_investigation_sid]"}              // earlier session_search investigation
]
FAST PATH — ~11 ms total
STEP 4F — FAST OUTPUT
Wrap snippets + context with metadata, return JSON
3,854 bytes (~963 tokens). No transcript fetch, no LLM, no truncation pass.
FAST — full response (Result 1 of 3)
{
  "success": true,
  "mode": "fast",
  "query": "[q1 topic keywords]",
  "results": [
    {
      "session_id": "cron_[sid_A]_20260501_1600",
      "when": "May 01, 2026 at 04:00 PM",
      "source": "cron",
      "model": "claude-opus-4-6",
      "matched_role": "assistant",
      "title": null,
      "snippet": "...[q1] [topic] [keywords] Task Created\\nThe user's **#1 priority for May 1, 2026** was established as a specific scoping task from the data pipeline:\\n- **Session ID:** `[q1-topic-sid]`\\n- **Goal:** apply a targeted transform to [topic] data...",
      "context": [
        {"role": "assistant", "content": "Now let me search for each session ID to find what was worked on today."},
        {"role": "tool", "content": "{\"success\": true, \"query\": \"[q1-topic-sid]\", \"results\": [{\"session_id\": \"cron_[sid_C]_20260501_1100\", \"when\": \"May 01, 2026 at 11:00 AM\", \"source\": \"cron\", \"model\": \"claude-opus-4-6\",...[trimmed]"},
        {"role": "tool", "content": "{\"success\": true, \"query\": \"session-search-investigation\", \"results\": [{\"session_id\": \"[earlier_investigation_sid]\", \"when\": \"April 23, 2026 at 10:34 AM\", \"source\": \"tui\", \"model\": \"claude-opus-4-6\", \"summ...[trimmed]"}
      ],
      "summary": "[Search hit — summary not generated in fast mode] Use snippet/context fields, or set mode='summary' for LLM-generated recall."
    }
  ],
  "count": 3, "sessions_searched": 3,
  "message": "Fast search returned FTS snippets without LLM summarization."
}
What the agent learns from fast mode
  • Three sessions touched the topic; here are their IDs and timestamps
  • Snippet shows one assistant turn mentioning the topic as Item #1
  • Context shows the agent was running session_search inside this cron job
  • NOT visible: the actual work session [work_session_sid] was never matched (cron jobs that mention the topic outranked it on FTS5 BM25)
SUMMARY PATH — ~36.7 s total
STEPS 4S–6S — FETCH, FORMAT, TRUNCATE, FAN OUT
db.get_messages_as_conversation × 3, then 100K-cap truncate, then 3 parallel auxiliary LLM calls
Per-session: ~25K input tokens fed to LLM with 30s timeout, 3 retries. The cost is here.
SUMMARY — system_prompt sent to auxiliary LLM
You are reviewing a past conversation transcript to help recall what happened.
Summarize the conversation with a focus on the search topic. Include:
1. What the user asked about or wanted to accomplish
2. What actions were taken and what the outcomes were
3. Key decisions, solutions found, or conclusions reached
4. Any specific commands, files, URLs, or technical details that were important
5. Anything left unresolved or notable

Be thorough but concise. Preserve specific details (commands, paths, error messages)
that would be useful to recall. Write in past tense as a factual recap.
SUMMARY — user_prompt template (filled per session)
Search topic: [q1 topic keywords]
Session source: cron
Session date: May 01, 2026 at 04:00 PM

CONVERSATION TRANSCRIPT:
{the ~100K-char truncated transcript from step 5S}

Summarize this conversation with focus on: [q1 topic keywords]
STEP 7S — SUMMARY OUTPUT
Wrap LLM outputs with metadata, return JSON
15,137 bytes (~3,784 tokens). 3.9× bigger than fast.
SUMMARY — full response (Result 1 of 3)
{
  "success": true,
  "mode": "summary",
  "query": "[q1 topic keywords]",
  "results": [
    {
      "session_id": "cron_[sid_A]_20260501_1600",
      "when": "May 01, 2026 at 04:00 PM",
      "source": "cron",
      "model": "claude-opus-4-6",
      "summary": "# Conversation Recap: [topic] Progress Check (May 01, 2026, 16:00)\n\n## Context\nAutomated cron check-in reviewing the day's planning doc. [topic] was Item #1 on the agenda.\n\n## Actions Taken\n1. Read the planning doc\n2. Ran `session_search` on the agenda session IDs\n3. Cross-referenced with recent sessions to find post-11:00 activity\n4. Updated the Progress section with a 16:00 entry\n\n## Key Findings\n**Status: \ud83d\udfe1 In progress — recon complete, implementation not started**\n\nThe morning work session (`[work_session_sid]`) progressed substantially beyond the 11:00 snapshot: discovery phase identified the relevant modules and the gaps in existing handling; full-dataset scan results shaped the planned approach… [truncated for length]"
    }
  ],
  "count": 3, "sessions_searched": 3
}
What the agent learns from summary mode
  • A multi-section structured recap: Context, Actions, Findings, Files, Outcome, Notable
  • References the morning work session [work_session_sid] (180 messages) — but reads it second-hand: the session itself was never matched by FTS5; the recap reproduces what the cron transcripts say about it
  • Reports specific findings: a measured count, dataset size, classification breakdown, and planned approach, status 🟡 in progress — accurate in this case, but every fact is downstream of the same cron transcripts the FTS5 hits returned
GUIDED PATH — fast → agent-think (~15s) → guided (~2ms)
STEP 4G — INPUT (composed: requires fast first)
Agent reads top fast hit, calls back with session_id + match_message_id
No FTS5 in this call (already paid for in the prior fast call), no auxiliary LLM, no truncation. Single db.get_messages_around() indexed range query.
GUIDED — input arguments
{
  "mode": "guided",
  "session_id": "cron_[sid_A]_20260501_1600",   // copied from fast result 1
  "around_message_id": 12793,                           // copied from fast result 1's match_message_id
  "window": 5
}
STEP 5G — DB FETCH
db.get_messages_around() — two indexed range queries
SQL: SELECT * FROM messages WHERE session_id=? AND id≤? ORDER BY id DESC LIMIT 6 (anchor + 5 before, reversed) UNION SELECT … AND id>? ORDER BY id ASC LIMIT 5 (5 after). Sorted ASC client-side. ~0.7 ms wall.
GUIDED — full response (49,371 bytes)
{
  "success": true,
  "mode": "guided",
  "session_id": "cron_[sid_A]_20260501_1600",
  "around_message_id": 12793,
  "window": 5,
  "session_meta": {
    "when": "May 01, 2026 at 04:00 PM",
    "source": "cron",
    "model": "claude-opus-4-6",
    "title": null
  },
  "messages": [
    {"id": 12788, "role": "assistant", "content": "...", "tool_calls": [...]},        // before-5
    {"id": 12789, "role": "tool",      "content": "{\"output\": \"20260501\", ...}"}, // before-4: date check
    {"id": 12790, "role": "assistant", "content": "", "tool_calls": [...]},          // before-3
    {"id": 12791, "role": "tool",      "content": "{\"content\": \"     1|# 2026-05-01 — Daily Plan (draft)\\n     2|\\n     3|## Carry-..."},  // before-2: read planning doc
    {"id": 12792, "role": "assistant", "content": "Now let me search for each session ID to find what was worked on today."},  // before-1
    {"id": 12793, "role": "tool",      "content": "{\"success\": true, \"query\": \"[q1-topic-sid]\", \"results\": [{\"session_id\":...}",
     "anchor": true},                                                                  // ← ANCHOR — session_search result for [q1-topic-sid] (the FTS5 match was on the keywords inside this tool result)
    {"id": 12794, "role": "tool",      "content": "{\"success\": true, \"query\": \"session-search-investigation\", \"results\": [...]}"}, // after-1
    {"id": 12795, "role": "tool",      "content": "{\"success\": true, \"query\": \"dashboard-sessions-order\", \"results\": [...]}"},   // after-2
    {"id": 12796, "role": "assistant", "content": "The session searches are mostly returning prior cron check-ins. Let me search fo..."},  // after-3
    {"id": 12797, "role": "tool",      "content": "{\"success\": true, \"query\": \"[topic] [keywords] recon scripts ...}"},  // after-4
    {"id": 12798, "role": "tool",      "content": "{\"success\": true, \"mode\": \"recent\", \"results\": [...]}"}              // after-5
  ],
  "messages_before": 5,
  "messages_after": 5
}
What the agent learns from guided mode
  • The 11 actual messages around the FTS5 anchor — verbatim, no LLM in the loop
  • Anchor message (id=12793) is itself a tool-result: the cron's previous session_search call for "[q1-topic-sid]". So this drill-down lets the agent see the tool result the original FTS5 match was scoring against — useful because the previous cron's session_search result contains the project status snapshot at that moment
  • The before-window shows the cron reading the planning doc (id=12791) and announcing its plan (id=12792). The after-window shows the cron continuing its own searches and pivoting (id=12796) when it sees only cron check-ins coming back
  • Caveat: guided still operates within the cron transcript that FTS5 retrieved — it doesn't reach the actual work session any more than fast or summary do. What it does give is the conversation around the match, which is more grounded than summary's cross-session reconstruction
  • Cost: guided's 49 kB payload is the full conversation window — bigger than summary's 12 kB recap. Composed across both turns: 53 kB of tokens enters the agent's context vs 12 kB for summary alone
⚠️ Critical: fast and summary share the same retrieval set

Both modes return the same three sessions. The mode parameter doesn't change which sessions are surfaced — only what's done with them. Fast wraps the FTS5 snippet and ±1 message context (~3 short messages per hit). Summary fetches each surviving session's full message list, truncates each to 100k chars, and dispatches parallel auxiliary LLM calls to recap them — see §6 for the per-query data flow.

For Q1 in particular: FTS5 returned three cron-job sessions that discuss the the Q1 topic — not the work session itself. Summary's recap is therefore a recap of cron transcripts that talk about the work, not a recap of the work. The chain is: actual work happened → cron job summarised it into its own transcript → FTS5 matched the cron transcript → our auxiliary LLM summarised the cron transcript → the calling agent reads that summary as if it were the source.

This isn't wrong per se — the cron summaries are themselves close to the truth — but it means summary mode's polished, confident output can hide retrieval gaps. The higher-leverage fix isn't between fast and summary: it's improving FTS5 retrieval so that both modes start from the right set of sessions.

5. Consumer simulation

What the user actually sees: replaying the agent's next turn

This section answers a different question from §4. §4 measured how long session_search takes to run (the tool-side latency where the mode choice actually changes the cost). §6 holds tool-side latency aside and asks: given the same JSON the live agent would receive, what answer does the model produce for the user? The replay calls the same auxiliary path session_search itself uses (auxiliary.session_search → auto → main chat model), feeding the JSON from the §4 run as the tool result. Same model, same payload, real LLM tokens.

⚠️ The mechanism: fast and summary differ in what each internally processes, not just what they emit

Both modes start with the same FTS5 hit set. After that, the paths diverge sharply:

  • Fast mode wraps each hit's snippet (~150 chars highlighted) plus ±1 message of context — typically 3 short messages per hit. For 3 sessions that's ~3.5–3.9 kB of JSON returned to the calling agent. No additional DB query, no auxiliary LLM call.
  • Summary mode takes the surviving sessions and, for each one, issues a second DB query (db.get_messages_as_conversation(session_id)) that pulls the entire session's messages. The full transcript is formatted, then truncated to 100,000 chars per session (centred on FTS5 match locations with a 25/75 recency bias). The (up to 100k-char) truncated transcripts are dispatched as parallel auxiliary LLM calls with a system prompt asking for a structured recap. Only the recap text — typically 1–4 kB per session — is returned.

So summary mode internally processes far more raw transcript than fast surfaces to the main loop, but the calling agent never sees that raw transcript — it sees a distillation. The numbers below quantify the gap.

Data flow per query — what each mode reads from the DB and sends to which LLM
Query sessions msgs fetched full chars → aux LLM (after 100k cap per session) cap hits main-loop bytes (fast) main-loop bytes (summary)
Q13127293k220k1 / 33.9k15.1k
Q2 investigation3413696k231k2 / 33.8k13.1k
Q3 planning3422526k200k2 / 33.5k10.1k
Q4 atlas0000—0.14k0.14k

Numbers from profile_session_search.py raw output, plus a follow-up query using the same DB snapshot. "msgs fetched" is the total across all surviving sessions for that query; "→ aux LLM" sums the per-session truncated transcripts that get dispatched in parallel; "main-loop bytes" is the JSON the calling agent sees, before the next-turn synthesis call.

Internal processing — summary reads ~57–62× more raw text

Summary mode's auxiliary LLM ingests ~200–230 kB of formatted transcript per query (across 3 parallel calls). Fast mode's user-visible JSON is ~3.5–3.9 kB per query.

That's where summary's 27–37 s wall time comes from: not a single big call, but three parallel calls each on ~70–100 kB of transcript, with the slowest dominating.

Truncation is real — H2 confirmed for Q2 and Q3

2 of 3 sessions in both Q2 and Q3 exceeded the 100k cap. The auxiliary's recap on those sessions is reading a window — 25% before / 75% after the first FTS5 match — not the whole transcript.

Q2's April 23 session is 397k chars, capped to 100k. Q3's April 23 session is the same. Late-session content (resolutions, decisions, status flips) is at risk of falling outside the window — a candidate follow-up beyond this PR.

Main-loop context cost — summary is ~3.4× fast (not 50×)

The 200 kB stays inside session_search — the calling agent only ever sees the recap text, around 10–15 kB total per query. So the headline tool-result inflation for the main loop is bounded at ~3.4×.

The cost the calling agent does pay is wall-clock time waiting for the auxiliary LLM (27–37 s) and a more polished but second-hand recap. The volume cost is absorbed by the auxiliary path.

Why this changes the framing

Before this measurement: "fast and summary share the same FTS5 set; summary's value-add is bounded by what an LLM can wring from the same sessions." That's literally true at the FTS5-hit layer.

But after the FTS5 hits: summary's auxiliary LLM also reads ~50× more transcript than fast surfaces. That extra material can include real signal the snippet+context didn't carry — and the recap that lands in the calling agent's context is also ~3.4× richer than fast's snippet payload. So the trade-off cuts both ways:

  • Fast: the calling agent sees the raw FTS5 snippets — small, honest about what was matched, but blind to anything outside the ±1-message window.
  • Summary: the calling agent sees a distillation drawn from up to 100k chars per session — richer, but at the cost of being second-hand and vulnerable to the truncation cliff (§5).

Neither is a clean win for "what the calling agent needs to know". Guided is the third lever in this trade-off — pull the actual messages around a chosen anchor, without summarisation and without a 100k-char truncation gamble.

Replay results — answer size & shape per (query, mode), with wider fast (limit=5) + multi-anchor guided (top-3)
Query (user-style phrasing)fast: charssummary: charsguided: charsverdict
Q1 — "what was the final design for the Q1 work?"1,6943,0793,432Summary identifies the four-pass span-manifest pipeline as final; guided flags its own retrieval gap explicitly
Q2 — "what is the session_search_investigation session?"1,8252,0712,709Summary names the root cause (_truncate_around_matches, 100k window); guided shows where investigation paused
Q3 — "find the latest planning session"1,3732,1382,871All three deliver; guided surfaces the noon pivot and 5-item agenda via multi-anchor
Q4 — "where did we land on atlas commons rewording?"747702—Still no FTS5 hits — wider limit didn't rescue the AND-default footgun

Fast and summary use limit=5; guided drills the top-3 fast hits in a single multi-anchor call. Quality assessment is qualitative. Latencies are tool-side from §4.

Q1 — wider hit list changed summary's conclusion

With 5 sessions in the recap window (vs 3), summary now names the later four-pass span-manifest pipeline on feat/[impl-branch] as the final design — explicitly contrasting it with an earlier standalone transform. A different conclusion from the 3-hit run; wider retrieval pulled in later cron transcripts that documented the pivot. Guided on the same query flags its own limit honestly: "the design above is the last fully documented design, but the implementation status beyond May 2 is not visible in these results." That's the epistemic honesty the page has been calling for.

Q2 / Q3 — multi-anchor surfaces cross-session structure

With 3 anchors per drill, guided reconstructs the full timeline (Q2) and the shape of the planning evolution (Q3) — neither came through on the prior single-anchor runs. Multi-anchor lets the agent see the arc across sessions, not just one cross-section.

Q4 — wider limit didn't rescue retrieval miss

FTS5 default-AND still returned 0 hits — bumping the limit from 3 to 5 doesn't help when there's nothing to rank. The cleanest evidence on the page that retrieval is the deeper lever: steering improvements (wider hit list, multi-anchor, leaner payloads) all need something to retrieve.

⚠️ The structural insight — retrieval is the deeper lever

Both modes work from the same FTS5 retrieval set. Summary's value-add is bounded by what an auxiliary LLM can wring from the same sessions the snippet path already saw. When retrieval gets the right sessions, both modes deliver useful answers (Q2, Q3); when retrieval misses (Q1: cron transcripts instead of the work session; Q4: 0 hits), summary's confident prose can hide the miss in a way fast's snippets do not. Polished output without uncertainty signalling is a liability when retrieval is imperfect.

This reframes the optimisation target. The mode parameter is optionality, not a ranking. The highest-leverage follow-up isn't between fast and summary — it's improving retrieval itself so every mode starts from the right sessions: (a) auto-OR rewriting for multi-keyword queries (currently default-AND footgun), (b) title and recency-aware ranking on top of FTS5 BM25, (c) potentially a semantic-search complement for "what was the design / decision / latest update" question shapes. Bigger than this PR — but the investigation's biggest takeaway for future work.

Replay methodology — pitfall noted

The first replay run produced 0-character completions for Q1-fast and Q3-fast — the model emitted finish_reason=stop with no content despite the JSON containing usable hits. This wasn't a real fast-mode failure: it was the harness's system prompt ("do not call any more tools") interacting with assistant-with-tool-call message scaffolding. Adding "even if partial, share what you can see" resolved it. A real agent in a real conversation has no equivalent constraint.

The corrected harness (replay_session_search_consumer.py) is reproducible alongside profile_session_search.py — harness pair available in the investigation workspace.

Follow-up measurement — three-run prompt-tuning arc on guided synthesis

A separate measurement on the same multi-anchor drill, run three times with progressively richer prompts, asking what changes the user prompt produces on the same retrieved messages. Recorded here as evidence rather than speculation — partial falsification is real data. All three on Opus, same DB, same query (session-search-investigation).

The underlying mechanism: where synthesis happens differs by mode. Summary pre-digests through the auxiliary LLM server-side (briefing prose is built in). Guided returns raw messages and the main agent does the synthesis from scratch — so the main agent's structural choices are entirely shaped by the user prompt. Guided is the most prompt-tunable mode.

ModeWho synthesisesWhat the user prompt can shape
summaryAuxiliary LLM (server-side)Limited — aux-prompt is fixed; your phrasing affects retrieval, not synthesis
fastMain agent, from snippetsLimited — little synthesisable material; mostly shapes how the agent navigates from there
guidedMain agent, from raw messagesHigh — synthesis structure is entirely the main agent's choice, prompt cues land directly

Run 1 — bare drill instruction. "drill into the most promising ones with multi-anchor guided"

  • Chronological state-walk with named stages, analyst-notes feel
  • Surfaced action items at the bottom (drift PR, related-PR status)
  • Some confabulation about PR state (mixed up #20238 with the salvage branch)
  • Honest prose, but not particularly polished — analyst's open notebook

Run 2 — added explicit briefing cue. "…give me a briefing-style recap, not a transcript walk"

  • Polished narrative arc with phase headings, story-shape closer
  • Verbatim quote pulling from past sessions — pulled the user's "fast session search that returns the wrong results is pretty useless" verbatim, unprompted
  • Lost the action items — narrative shape dominated, operational handoff dropped
  • Falsified the original "guided defaults to transcript walk" hypothesis — both runs produced prose

Run 3 — full stack: briefing + tightness + active retrieval + dual closer. The full prompt is reproduced in the investigation workspace alongside the harness pair.

  • Active retrieval played hard: initial fast → multi-anchor guided → refetch with smaller windows → second supplemental fast search ("guided mode session_search implementation OR merged OR PR") to plug a known shipping-state gap. The model didn't synthesise from the original hits — it noticed a gap and went and got more data.
  • Briefing prose survived the workflow instructions. Output stayed narrative-shaped — Scope, four numbered Stages, Handoff, Gaps. Tightness cue held: bulleted lists where they should be, paragraphs where they should be, no sprawl.
  • Both closer sections present and substantive. Handoff = 6 actionable bullets. Gaps = 5 bullets, each pointing at a real retrieval failure (meta-pollution, missing implementation session, thin Stage 1 anchor, no guided-mode metrics) — not hedging.
  • Verbatim quoting continued unprompted — appears to be what Opus does once it has raw guided messages with user dialogue; no explicit cue needed.
  • The killer self-aware bullet: the gaps section ended with — "Live demonstration of the bug: ironically, this whole briefing is itself an example of the pathology — I'm reconstructing the Stage 4 shipping state from cron/sibling references rather than from the work session, because that's what FTS5 surfaced. Exactly the failure mode the investigation documented." The agent caught itself instantiating the very bug being investigated.

Refined finding: the steering lever isn't transcript-vs-prose — it's how much structural opinion the agent invents vs how much the user specifies. Four shape-knobs mattered for guided synthesis: briefing framing, tightness, active retrieval, and dual closer (handoff + gaps as explicit named sections). All four landed in Run 3.

Why this matters beyond styling: Run 3's gaps section surfaced a finding (the meta-pollution self-demo) that wouldn't have been visible from Run 1 or Run 2. The honest-about-gaps cue isn't just aesthetic — it's an epistemic surface that catches when retrieval is quietly failing. Same query, same modes, same data; just better instructions to the synthesiser.

Tool-side speedup isn't free at the agent-context level

A measurement that complicates the headline "fast is faster than summary": guided is never used standalone — it always rides on a prior fast call plus a main-agent inference turn (~15 s on Opus) that picks the anchors. The honest end-to-end cost of "fast → guided" as a drill-down workflow is in the same order of magnitude as summary's tool wall, dominated by inference rather than the sub-millisecond tool walls themselves.

And the payload trade-off cuts the other way too: on Q1, multi-anchor guided returned ~131 kB to the calling agent's context (three windows of raw messages, some with multi-kB tool outputs) vs summary's 21 kB recap. The composed payload across both turns is fast (~6.6 kB) + guided (~131 kB) = ~138 kB, vs summary's single 21 kB.

Lesson: tool-side speedup isn't free at the agent-context level, especially when the "cheap" path is two tool calls deep. Quality-wise guided still wins on Q1 (it surfaced the four-pass span-manifest design that summary missed) — but speed is not the reason to reach for it.