Session Search — Optimisations
value-prop and measured outcomes — May 2026
You used to have one way to ask the system about its memory: a 30–60 second LLM-narrated recap. You now have three — a cheap snippet scan, a deep recap, and a raw-transcript drill-down — and the agent picks the right one for the question.
One mode for everything → three modes the agent picks between
session_search had one mode. Every question — discovery, drill-down, synthesis — paid the same ~30–60s, ~$1.00 for an LLM-narrated recap. Worse: when retrieval was weak, the aux LLM still produced confident prose. Thin hits got laundered into authoritative-sounding answers.
Guided is the deep-dive after fast or summary has surfaced where to look.
What actually runs, when, for how long
Same question, three modes, time flowing left-to-right. Each track shows where the wall-clock seconds and dollars go — agent reasoning (grey), tool calls (coloured), aux LLM work (red).
The summary track is dominated by aux-LLM work — that's where the cost lives. Other modes are tool reads against SQLite; the visible time is mostly the agent reasoning between calls and writing the final reply.
Three shapes of question, three behaviours, no flags to juggle
The agent reads the shape of your question and picks. You can override — name a mode in your prompt and it'll be honoured — but for most usage the defaults land correctly.
auxiliary.session_search.default_mode: fast
in ~/.hermes/config.yaml. The agent will start with cheap snippet scans on most questions
and only reach for synthesis when the question explicitly demands it.
There's no pre-routing step, no classifier, no separate model deciding which mode to use. Mode selection is the main model's own choice, made every turn from three signals — all of which live in the tool's schema description, the shipped session-recall skill, and the user's request:
- Schema description — the tool description is a compact specification: what each mode does, returns, and costs; the anchor contract; FTS5 syntax; and a one-paragraph "when to use". It ships in every system prompt and stays tight.
session-recallskill — the playbook layer. Pre-flight rules, mode picker table, levers, composition patterns, worked examples, pitfalls. Auto-loaded by the agent on recall-shaped questions; not in every system prompt, so cost is paid only when relevant.- User wording — phrases like "drill into…", "walk me through…", "find the session where…", or an explicit "use fast mode" shift which signature the model matches. Neither schema nor skill enumerates these phrases; the model generalises from the role tags (discovery / synthesis / drill-down).
- Configured default — when the model omits
mode, the resolver readsauxiliary.session_search.default_mode(defaultfastas of this PR). Explicitmode=always wins over the configured default.
The implication: mode-selection quality is a prompt-engineering surface, not a code surface. The schema specifies what; the skill teaches when and why. Both can be sharpened independently of the implementation.
Three modes return three different shapes of data
Same DB, same query, different artefacts handed to the calling agent. Wide-and-shallow snippets, narrative prose, or raw conversation windows.
SNIPPETS FTS5 lifts a few sentences around each match. Just enough to recognise which session is the right one to drill into.
analytics.py, adding a regression test, and merging via PR #20183. Key decisions: …"NARRATIVE PROSE Aux LLM reads each matched session top-to-bottom and writes a recap. Cost scales linearly per matched session.
RAW CONVERSATION + BOOKENDS Verbatim messages around the anchor, plus the first/last few user+assistant turns so you get the session's framing. Tool messages dropped except the anchor itself. No LLM rewriting.
~1 KB per session
~4 KB per session
up to ~130 KB
Fast and summary are horizontal — they scan across sessions. Guided is vertical — it dives into one. The typical flow is fast (find the candidates) → guided (read what was actually said). Summary collapses both into a single LLM-narrated middle, at LLM-narrated cost.
The fast → guided composition is only as good as the anchor fast picks. If match_message_id lands on a tangential snippet — a passing mention, a meta-discussion, the wrong session of a multi-session topic — guided will dutifully fetch the ±N messages around the wrong point. Three mitigations are built in:
- Bookends partially compensate. Even when the window is centred on noise,
bookend_startandbookend_endsurface the session's actual opening and resolution. The reader at least gets the session's shape for free, regardless of where the anchor landed. - Multi-anchor fan-out for ambiguous topics. When a subject spans multiple sessions, the schema teaches the agent to pass the top 2–3 fast hits as separate anchors in one guided call — so a bad top hit doesn't sink the whole drill. This is the design lever for the "FTS5 ranks by relevance, not recency" failure mode (the top hit for a multi-day arc is often the oldest session).
- Re-fast with sharper query. When guided windows look noisy or empty, the agent can re-run fast with refined wording. This is unguarded — relies entirely on the model noticing — and is the most fragile of the three.
What's not built in: automated re-anchoring, anchor-quality scoring, or a "this looked wrong, try again" loop. Anchor selection failure is a real failure mode this spike doesn't fully solve — the bookends and multi-anchor fan-out close most of the gap, but a determined adversarial query could still produce a confidently-centred window of irrelevant context. Worth measuring in a follow-up.
(session_id, message_id) pairs — typically picked from a prior fast call's results.user,assistant (tool noise filtered). Set to user,assistant,tool to include tool output.default_mode config.
The agent composes these knobs without you having to think about them. The shape of your question drives mode; the breadth of the topic drives limit; the depth of the drill drives anchors and window. Power users can pin any of them — the agent honours explicit overrides.
What "honest" means now
The biggest qualitative change isn't speed or cost — it's that the system stops papering over weak retrieval with confident-sounding prose.
tone
tone
The old default (summary-on-everything) sat on the dangerous diagonal. The new defaults pull toward "honest about gaps" by surfacing raw snippets first — and let the agent widen the search when results look thin.
Aux LLM summarises whatever FTS5 found. If retrieval missed the right session, the reader gets a confident paragraph built from adjacent noise. No signal that retrieval was thin.
Agent runs fast, surfaces 2 telegram sessions with snippets, calls guided on the most substantive one for raw messages. Synthesises from actual conversation: founders, manifesto URL, the @telepathinc tease.
Bonus on prior-work questions: the agent now reaches for the session DB before gh pr list / web / filesystem. Memory first, external sources second.
One anecdote is a story; fourteen is a signal. Below: each smoke-test scenario, the mode the agent actually picked, the retrieval strength observed, and the specific risk the old summary-on-everything default would have run.
93% of scenarios (14/15 excluding the 3 test-infrastructure cases) would have run measurably worse under the old default — either by paying multiples of the new cost, by silently laundering thin retrieval into confident prose, or by missing a behavioural improvement (memory-first instinct, temporal sorting cue, multi-anchor catch-up). v2 numbers reflect 90 iterations across 18 scenarios; cell verdicts are median behaviour over 5 iterations each.
For each scenario: min → max range of wall time across 5 iterations, with median (teal) and p90 (yellow) ticks. Most scenarios cluster tight; S05 spreads wide (regression test exercising the pair-resolution safety net), and S14 is the highest absolute (premium-prompt active retrieval).
Same shape, in dollars. The cheap scenarios (S06, S11, S12) show one-outlier sensitivity — single $0.18 iteration in a sea of $0.03 iterations spreads the range; the median holds steady. Synthesis scenarios cluster tight around $0.30; S14 is the cost ceiling at ~$0.65.
Each scenario judged on Mode appropriateness, Search fidelity, Topic fidelity, Usefulness — 1–5 scale, grader independent of the run agent. Predominantly green; the 7 with score docks are annotated.
Where the dollars actually go
Costs are split between the main agent loop (your conversation with Opus) and the auxiliary LLM work that summary mode dispatches per matched session. Both are priced from the same rate card.
v2 numbers: medians and p90s computed across 5 iterations × 18 scenarios. Pricing computed by hermes's own agent/usage_pricing.py against the Anthropic claude-opus-4-7 rate card (May 2026) — same rate card for both main and aux costs. Ratios are robust; absolute dollars assume Opus pricing and will scale with provider/model choice.
The summary row carries its v1.8 reference value because v2's agent didn't pick summary in any of the 90 iterations — even on S13 where default_mode: summary was configured (see §3 callout on the schema-teaching ceiling). The row is preserved here so the reader can see the cost trade-off summary represents; under the new schema it's effectively opt-in only.
browse isn't one of the three modes — it's session_search() called with no query, a separate code path that lists recent sessions. Included here as a baseline because the agent reaches for it on "what was I working on recently?"-shaped questions.
Three takeaways, one config line
What this means in practice
- The default is now fast. Schema reframed so "catch me up on X" routes through
fast → guidedinstead of reflexively burning a summary call. Same DB, same retrieval, but recall now starts cheap. v2 measured this across 90 iterations on 18 scenarios — 17/18 showed 5/5 mode-pick stability. - Memory first, external sources second. The agent now reaches for the session DB before
gh pr list, web search, or the filesystem on questions about prior work. The behavioural shift is bigger than the cost shift — and the cost shift is real. - Guided is the recall workhorse. Returns raw messages around your anchor plus the session's opening and closing turns as bookends — so you get framing for free, without an aux LLM. Tool output is filtered out by default.
- Summary is opt-in for genuine synthesis. When you actually want a recap across N sessions, it's still there — but it's a deliberate choice now, not the default for every question that mentions the past.
- A shipped
session-recallskill teaches the composition. Every install bundles a skill atskills/memory/session-recall/that the agent loads on recall-shaped questions: pre-flight rules, mode picker, levers, composition patterns, worked examples, pitfalls. The schema describes what each mode does; the skill teaches when and why. - Honesty improved more than cost. No hallucinations observed across the 90-iteration v2 set. Rubric scoring: 11/18 scenarios perfect 5/5/5/5, 7 with small documented score docks (none disqualifying). Thin retrieval shows up as thin output instead of confident-but-wrong prose.
Two ways, depending on how much control you want.
# ~/.hermes/config.yaml — flip the default to fast
auxiliary:
session_search:
default_mode: fast # or "summary"
# Restart hermes, then use normally. The agent will start with fast on
# most questions and only reach for synthesis when you explicitly ask.
Or force a mode per call by naming it in the prompt: "use fast mode to find sessions about X",
"summary only", "drill into the most promising hit with guided". Explicit overrides
always win — both at the resolver level (the agent's explicit mode= argument bypasses the configured default)
and at the prompt level (the agent reads what you ask for and honours it).
Caveat measured in v2: setting default_mode: summary in config is currently a soft preference. The schema description tells the agent about your configured default, but on synthesis-shaped questions ("recall…", "walk me through…", "catch me up on…") the agent's training that fast → guided is the right composition currently overrides the user's stated preference. If you want guaranteed summary mode, name it in the prompt ("use summary mode for this") — that always works.
Read the deep-dive →