Session Search — Optimisations

value-prop and measured outcomes — May 2026

You used to have one way to ask the system about its memory: a 30–60 second LLM-narrated recap. You now have three — a cheap snippet scan, a deep recap, and a raw-transcript drill-down — and the agent picks the right one for the question.

THE PROBLEM

One mode for everything → three modes the agent picks between

session_search had one mode. Every question — discovery, drill-down, synthesis — paid the same ~30–60s, ~$1.00 for an LLM-narrated recap. Worse: when retrieval was weak, the aux LLM still produced confident prose. Thin hits got laundered into authoritative-sounding answers.

The three modes
fast
starting move · discovery FTS5 snippets and metadata across N candidate sessions. No LLM. ~1 KB per session, ~14 seconds end-to-end.
summary
starting move · synthesis An auxiliary LLM reads each matched session and writes a recap. ~4 KB per session, ~45–60 seconds. The previous default — now opt-in.
guided
follow-up move · drill-down Raw messages around a chosen anchor, plus the first/last few user/assistant messages as bookends. No LLM, no rewriting. Requires an anchor from a prior fast or summary call.
What each mode actually does
fast
STEP 1 · FTS5 Query the full-text index for matches (user + assistant roles by default; tool noise filtered)
RETURN Snippets + metadata for each match
cost: index read only
summary
STEP 1 · FTS5 Same query against the index
STEP 2 · FETCH Pull all messages from every matched session
STEP 3 · AUX LLM Summarise each session's transcript with Opus, in parallel
RETURN One LLM-written recap per matched session
cost: 1 aux LLM call per matched session
guided
STEP 1 · FETCH For each anchor, pull ±N messages around it (tool noise filtered)
STEP 2 · BOOKENDS Also pull the first & last few user/assistant messages of the session for framing
RETURN Raw conversation window + bookends per anchor
cost: targeted DB read only
The cost gap isn't about the search. Fast and summary run the same FTS5 query. The gap is what happens after: summary fetches every message from every match and runs an aux LLM over each one. Fast skips both. Guided skips the search entirely and just reads a window around a known point, plus the session's framing bookends.
How they connect
two starting moves
question
→
fast
question
→
summary
one follow-up move
fast
→
guided
summary
→
guided

Guided is the deep-dive after fast or summary has surfaced where to look.

ON ONE CLOCK

What actually runs, when, for how long

Same question, three modes, time flowing left-to-right. Each track shows where the wall-clock seconds and dollars go — agent reasoning (grey), tool calls (coloured), aux LLM work (red).

Median wall time and total cost, by mode pattern
browseno query
agent replies
$0.15~11s
fast onlydiscovery
think
agent replies
$0.17~14s
fast + guideddrill-down
think
picks anchor
agent replies
$0.28~28s
fast×N + guidedactive retrieval
think
sharpen
sharpen
sharpen
sharpen
sharpen
picks anchors
·
agent replies (briefing)
$0.69~72s
summarysynthesis
think
agent replies
$1.00~60s
agent reasoning fast (snippets) guided (drill-down) browse (no query) aux LLM (summary work)

The summary track is dominated by aux-LLM work — that's where the cost lives. Other modes are tool reads against SQLite; the visible time is mostly the agent reasoning between calls and writing the final reply.

QUESTION → MODE

Three shapes of question, three behaviours, no flags to juggle

The agent reads the shape of your question and picks. You can override — name a mode in your prompt and it'll be honoured — but for most usage the defaults land correctly.

Question shape
Mode picked
Typical wall
Typical cost
"find the session where we…" "which conversation discussed the dashboard cache-tokens fix?"
fast
~14s
~$0.17
"show me what we actually said when…" "drill into the May 12 session where we pivoted the design"
fast → guided
~28s
~$0.28
"what did we decide about…" "walk me through how the atlas-drift detector cron came together"
summary
~45–60s
~$1.00
The lever a power user wants: set auxiliary.session_search.default_mode: fast in ~/.hermes/config.yaml. The agent will start with cheap snippet scans on most questions and only reach for synthesis when the question explicitly demands it.
How mode selection actually happens

There's no pre-routing step, no classifier, no separate model deciding which mode to use. Mode selection is the main model's own choice, made every turn from three signals — all of which live in the tool's schema description, the shipped session-recall skill, and the user's request:

  • Schema description — the tool description is a compact specification: what each mode does, returns, and costs; the anchor contract; FTS5 syntax; and a one-paragraph "when to use". It ships in every system prompt and stays tight.
  • session-recall skill — the playbook layer. Pre-flight rules, mode picker table, levers, composition patterns, worked examples, pitfalls. Auto-loaded by the agent on recall-shaped questions; not in every system prompt, so cost is paid only when relevant.
  • User wording — phrases like "drill into…", "walk me through…", "find the session where…", or an explicit "use fast mode" shift which signature the model matches. Neither schema nor skill enumerates these phrases; the model generalises from the role tags (discovery / synthesis / drill-down).
  • Configured default — when the model omits mode, the resolver reads auxiliary.session_search.default_mode (default fast as of this PR). Explicit mode= always wins over the configured default.

The implication: mode-selection quality is a prompt-engineering surface, not a code surface. The schema specifies what; the skill teaches when and why. Both can be sharpened independently of the implementation.

WHAT COMES BACK

Three modes return three different shapes of data

Same DB, same query, different artefacts handed to the calling agent. Wide-and-shallow snippets, narrative prose, or raw conversation windows.

Per-call response anatomy (real bytes from the smoke test)
fast
~4 KB · 3 sessions
session_id, title, timestamp
"…dashboard cache-tokens fix landed in…"
match_message_id, source
session_id, title, timestamp
"…investigated cache-tokens in May 1 cron…"
match_message_id, source
+ 1 more match

SNIPPETS FTS5 lifts a few sentences around each match. Just enough to recognise which session is the right one to drill into.

summary
~12 KB · 3 sessions
session_id, title, timestamp
"The dashboard analytics fix involved patching the cache-tokens SQL query in analytics.py, adding a regression test, and merging via PR #20183. Key decisions: …"
session_id, title, timestamp
"The investigation traced back to a May 1 cron run where cache token attribution was lost in the rollup. The team isolated the bug to …"
+ 1 more, all LLM-generated

NARRATIVE PROSE Aux LLM reads each matched session top-to-bottom and writes a recap. Cost scales linearly per matched session.

guided
~12–130 KB · 1 anchor
session_id, anchor_message_id
bookend_start ← session opener
[user] "let's profile the FTS5 path…"
[assistant] "let me check the schema…"
⚓ [user] "but the truncation is dishonest"
[assistant] "you're right, I'll…"
bookend_end ← session closer
[assistant] "shipped 76f40e644 — ready"
tool noise filtered out

RAW CONVERSATION + BOOKENDS Verbatim messages around the anchor, plus the first/last few user+assistant turns so you get the session's framing. Tool messages dropped except the anchor itself. No LLM rewriting.

Where each mode sits on the breadth ↔ depth axis
BREADTH many sessions, shallow context per session
fast
N sessions × snippet
~1 KB per session
summary
N sessions × LLM recap
~4 KB per session
guided
1 anchor × full window
up to ~130 KB
DEPTH one anchor, full surrounding context

Fast and summary are horizontal — they scan across sessions. Guided is vertical — it dives into one. The typical flow is fast (find the candidates) → guided (read what was actually said). Summary collapses both into a single LLM-narrated middle, at LLM-narrated cost.

When anchor selection goes wrong

The fast → guided composition is only as good as the anchor fast picks. If match_message_id lands on a tangential snippet — a passing mention, a meta-discussion, the wrong session of a multi-session topic — guided will dutifully fetch the ±N messages around the wrong point. Three mitigations are built in:

  • Bookends partially compensate. Even when the window is centred on noise, bookend_start and bookend_end surface the session's actual opening and resolution. The reader at least gets the session's shape for free, regardless of where the anchor landed.
  • Multi-anchor fan-out for ambiguous topics. When a subject spans multiple sessions, the schema teaches the agent to pass the top 2–3 fast hits as separate anchors in one guided call — so a bad top hit doesn't sink the whole drill. This is the design lever for the "FTS5 ranks by relevance, not recency" failure mode (the top hit for a multi-day arc is often the oldest session).
  • Re-fast with sharper query. When guided windows look noisy or empty, the agent can re-run fast with refined wording. This is unguarded — relies entirely on the model noticing — and is the most fragile of the three.

What's not built in: automated re-anchoring, anchor-quality scoring, or a "this looked wrong, try again" loop. Anchor selection failure is a real failure mode this spike doesn't fully solve — the bookends and multi-anchor fan-out close most of the gap, but a determined adversarial query could still produce a confidently-centred window of irrelevant context. Worth measuring in a follow-up.

The agent's tunable parameters
query
fast, summary
Keywords, phrases, boolean (OR/NOT), prefix wildcards. Omit to browse recent sessions.
limit
fast, summary
How many sessions to return. Default 3, max 10. Bump higher for wider breadth.
anchors
guided
One or more (session_id, message_id) pairs — typically picked from a prior fast call's results.
window
guided
Messages on each side of each anchor. Default 5, max 20. Larger = more context per drill.
role_filter
fast, summary
Restrict to specific roles. Defaults to user,assistant (tool noise filtered). Set to user,assistant,tool to include tool output.
mode
all
Explicit override. Honored regardless of default_mode config.

The agent composes these knobs without you having to think about them. The shape of your question drives mode; the breadth of the topic drives limit; the depth of the drill drives anchors and window. Power users can pin any of them — the agent honours explicit overrides.

SEARCH FIDELITY

What "honest" means now

The biggest qualitative change isn't speed or cost — it's that the system stops papering over weak retrieval with confident-sounding prose.

The failure mode the old version had
weak retrieval
strong retrieval
confident
tone
Confident-stale (the old failure) Summary mode laundering FTS5 misses into authoritative prose. The reader can't tell the answer is fabricated.
Correct and confident What you'd want every time. The old system could deliver this — when retrieval happened to work.
hedged
tone
Honest about gaps (the new default) Fast surfaces what it found and flags what's thin. Agent can re-search with sharper phrasing before synthesising.
Over-cautious Rare. The agent hedges even on strong retrieval — minor cost, easy to dial in via prompt.

The old default (summary-on-everything) sat on the dangerous diagonal. The new defaults pull toward "honest about gaps" by surfacing raw snippets first — and let the agent widen the search when results look thin.

A real example — same question, both modes
old default — summary on everything
"what do we know about Telepath?"
~45–60s, ~$1.00.

Aux LLM summarises whatever FTS5 found. If retrieval missed the right session, the reader gets a confident paragraph built from adjacent noise. No signal that retrieval was thin.
new default — fast, drill if needed
"what do we know about Telepath?"
~22s, ~$0.23.

Agent runs fast, surfaces 2 telegram sessions with snippets, calls guided on the most substantive one for raw messages. Synthesises from actual conversation: founders, manifesto URL, the @telepathinc tease.

Bonus on prior-work questions: the agent now reaches for the session DB before gh pr list / web / filesystem. Memory first, external sources second.

The whole smoke set, classified — would the old default have laundered?

One anecdote is a story; fourteen is a signal. Below: each smoke-test scenario, the mode the agent actually picked, the retrieval strength observed, and the specific risk the old summary-on-everything default would have run.

Scenario
Question shape
Mode picked
Retrieval
Old-default risk
S01
"what was I working on recently?"
browse
n/a (no FTS)
none — browse handles this
S02
"recall the atlas-drift cron arc"
fast → guided
strong (3 anchors)
~3× more cost for same answer
S03
"walk me through the cache-tokens fix"
fast → guided
strong (2 anchors)
would have paid summary cost
S04
"status of commons-messaging PR?"
fast (sort=newest) → guided
strong (5/5 sort=newest)
would have hit gh first, missed memory
S05
regression test — multi-anchor drill
fast + guided (60% rebind)
strong (safety net firing)
S06
"hermes-agent-dev skill?"
fast (alone)
strong (5/5 5/5)
~6× more cost for same answer
S07
"recap memory/context architecture"
fast → guided
strong (2-3 anchors)
would have paid summary cost
S08
"drill into X cold" (no prior fast)
fast → guided
strong (correct refusal/recovery)
would have laundered with summary
S09
"Karpathy llm-wiki ingestion?"
fast (sometimes sort=oldest)
strong (rare-token recall)
~6× more cost for same answer
S10
"atlas drift PR review"
fast (alone)
strong (no drill needed)
~6× more cost for same answer
S11
nonexistent topic (zero-hit probe)
fast
thin (no hits, $0.03)
the canonical confident-stale failure mode
S12
FTS5 default-AND footgun probe
fast
strong, $0.04
~6× more cost for same answer
S13
A/B mirror of S02 with default_mode:summary
fast → guided (config ignored)
strong but didn't honour config
S14
Run 3 prompt — active retrieval
fast×N + guided (sort=newest 5/5)
strong (briefing + handoff + gaps)
would not have fanned out / refined
S15
"where did we leave X" (recency-shaped)
fast (sort=newest 5/5) → guided
strong (1 anchor, $0.12)
old default had no temporal signal
S16
"how did X start" (origin-shaped)
fast (sort=oldest 5/5) → guided
strong (1-3 anchors)
origin hidden under descendants
S17
"catch me up across N sessions"
fast → guided (3 anchors 5/5)
strong (bookends + multi-anchor)
single-anchor would have missed half the arc
S18
anchor-quality stress test
fast → guided (sensible anchor 5/5)
strong (canonical moment found)
10
HIGH-risk under old default
Would have paid 6× for the same answer, missed a behavioural improvement (memory-first / temporal cue / multi-anchor), or laundered confident prose over a zero-hit retrieval.
4
MEDIUM-risk under old default
Same answer eventually, but billed at summary cost ($0.30+) instead of fast cost ($0.05–$0.20).
1
No change — old default was fine
S01 browse path is unchanged across versions.

93% of scenarios (14/15 excluding the 3 test-infrastructure cases) would have run measurably worse under the old default — either by paying multiples of the new cost, by silently laundering thin retrieval into confident prose, or by missing a behavioural improvement (memory-first instinct, temporal sorting cue, multi-anchor catch-up). v2 numbers reflect 90 iterations across 18 scenarios; cell verdicts are median behaviour over 5 iterations each.

Variance across 5 iterations — wall time

For each scenario: min → max range of wall time across 5 iterations, with median (teal) and p90 (yellow) ticks. Most scenarios cluster tight; S05 spreads wide (regression test exercising the pair-resolution safety net), and S14 is the highest absolute (premium-prompt active retrieval).

range median p90 · scale 0–90s
S01
10.1s
S02
38.1s
S03
28.9s
S04
22.9s
S05
28.2s
S06
9.9s
S07
47.3s
S08
27.6s
S09
16.0s
S10
20.5s
S11
9.5s
S12
12.5s
S13
38.6s
S14
66.9s
S15
17.8s
S16
33.7s
S17
28.2s
S18
32.0s
Variance across 5 iterations — cost

Same shape, in dollars. The cheap scenarios (S06, S11, S12) show one-outlier sensitivity — single $0.18 iteration in a sea of $0.03 iterations spreads the range; the median holds steady. Synthesis scenarios cluster tight around $0.30; S14 is the cost ceiling at ~$0.65.

range median p90 · scale $0–$0.75
S01
$0.158
S02
$0.354
S03
$0.310
S04
$0.156
S05
$0.222
S06
$0.049
S07
$0.342
S08
$0.379
S09
$0.185
S10
$0.097
S11
$0.030
S12
$0.036
S13
$0.355
S14
$0.649
S15
$0.118
S16
$0.333
S17
$0.273
S18
$0.296
Rubric scores — 4 axes × 18 scenarios (median across 5 iterations)

Each scenario judged on Mode appropriateness, Search fidelity, Topic fidelity, Usefulness — 1–5 scale, grader independent of the run agent. Predominantly green; the 7 with score docks are annotated.

11/18 scenarios perfect 5/5/5/5 across all four axes. 7 with at least one axis docked, all small, all explained.
5 · ideal   4 · minor nit   3 · noted issue   2 · meaningful dock  
S01
M5
S4
T5
U5
Rock-solid browse behaviour on the carry-over prompt; correctly identifies the active session and offers a guided next s
S02
M5
S5
T5
U5
S03
M5
S5
T5
U5
S04
M3
S5
T5
U5
Answers are excellent and consistent, but every run escalates to guided when fast alone was expected — systematic over-s
S05
M2
S4
T3
U3
Regression test split the agent: half drilled correctly, half flagged the fabricated SHA and refused — meta-honesty vs.
S06
M5
S5
T5
U5
S07
M5
S5
T5
U5
S08
M5
S5
T5
U5
S09
M5
S4
T5
U4
Fast mode reliably surfaces the rare-token Apr 17 sessions; variance is in whether the 18:02 companion gets directly hit
S10
M5
S5
T5
U5
S11
M5
S5
T5
U5
S12
M5
S5
T5
U5
S13
M3
S5
T5
U5
Answers are excellent and consistent, but default_mode=summary had zero observable effect — agent never used summary mod
S14
M5
S5
T5
U5
S15
M4
S5
T5
U5
Recency-shaped query reliably triggers sort=newest and anchors on the latest matching session; only nit is consistent gu
S16
M4
S5
T5
U5
sort=oldest reliably steers the agent to the originating session; quality of the origin framing varies but is consistent
S17
M5
S5
T5
U5
S18
M5
S5
T5
U5
COSTS

Where the dollars actually go

Costs are split between the main agent loop (your conversation with Opus) and the auxiliary LLM work that summary mode dispatches per matched session. Both are priced from the same rate card.

Median and p90 per-scenario cost, stacked by source — v2 (5 iterations × 18 scenarios)
browseno query
$0.16
$0.16 · p90 $0.16
fast onlydiscovery
$0.07
$0.07 · p90 $0.18
fast + guideddrill-down
$0.32
$0.32 · p90 $0.34
fast×N + guidedactive retrieval (Run 3 stack)
$0.65
$0.65 · p90 $0.66
summaryaux LLM per matched session (v1.8 reference)
main $0.49
aux LLM $0.52
$1.00 · n/a in v2

v2 numbers: medians and p90s computed across 5 iterations × 18 scenarios. Pricing computed by hermes's own agent/usage_pricing.py against the Anthropic claude-opus-4-7 rate card (May 2026) — same rate card for both main and aux costs. Ratios are robust; absolute dollars assume Opus pricing and will scale with provider/model choice.

The summary row carries its v1.8 reference value because v2's agent didn't pick summary in any of the 90 iterations — even on S13 where default_mode: summary was configured (see §3 callout on the schema-teaching ceiling). The row is preserved here so the reader can see the cost trade-off summary represents; under the new schema it's effectively opt-in only.

browse isn't one of the three modes — it's session_search() called with no query, a separate code path that lists recent sessions. Included here as a baseline because the agent reaches for it on "what was I working on recently?"-shaped questions.

The headline: summary is the only expensive mode, and the expense is the aux LLM call dispatched per matched session. Reserve it for synthesis questions — for everything else, fast is ~6× cheaper and ~4× faster ($0.17 vs $1.00, 14s vs 60s on these scenarios). Ratios are robust across providers; absolute dollar figures assume claude-opus-4-7 (Anthropic public rate card, May 2026). Swap providers and the dollar columns scale, the column-to-column ratios don't.
BOTTOM LINE

Three takeaways, one config line

What this means in practice

  • The default is now fast. Schema reframed so "catch me up on X" routes through fast → guided instead of reflexively burning a summary call. Same DB, same retrieval, but recall now starts cheap. v2 measured this across 90 iterations on 18 scenarios — 17/18 showed 5/5 mode-pick stability.
  • Memory first, external sources second. The agent now reaches for the session DB before gh pr list, web search, or the filesystem on questions about prior work. The behavioural shift is bigger than the cost shift — and the cost shift is real.
  • Guided is the recall workhorse. Returns raw messages around your anchor plus the session's opening and closing turns as bookends — so you get framing for free, without an aux LLM. Tool output is filtered out by default.
  • Summary is opt-in for genuine synthesis. When you actually want a recap across N sessions, it's still there — but it's a deliberate choice now, not the default for every question that mentions the past.
  • A shipped session-recall skill teaches the composition. Every install bundles a skill at skills/memory/session-recall/ that the agent loads on recall-shaped questions: pre-flight rules, mode picker, levers, composition patterns, worked examples, pitfalls. The schema describes what each mode does; the skill teaches when and why.
  • Honesty improved more than cost. No hallucinations observed across the 90-iteration v2 set. Rubric scoring: 11/18 scenarios perfect 5/5/5/5, 7 with small documented score docks (none disqualifying). Thin retrieval shows up as thin output instead of confident-but-wrong prose.
How to try it

Two ways, depending on how much control you want.

# ~/.hermes/config.yaml — flip the default to fast
auxiliary:
  session_search:
    default_mode: fast    # or "summary"

# Restart hermes, then use normally. The agent will start with fast on
# most questions and only reach for synthesis when you explicitly ask.

Or force a mode per call by naming it in the prompt: "use fast mode to find sessions about X", "summary only", "drill into the most promising hit with guided". Explicit overrides always win — both at the resolver level (the agent's explicit mode= argument bypasses the configured default) and at the prompt level (the agent reads what you ask for and honours it).

Caveat measured in v2: setting default_mode: summary in config is currently a soft preference. The schema description tells the agent about your configured default, but on synthesis-shaped questions ("recall…", "walk me through…", "catch me up on…") the agent's training that fast → guided is the right composition currently overrides the user's stated preference. If you want guaranteed summary mode, name it in the prompt ("use summary mode for this") — that always works.