Skip to content
← all work

Netision · 2025–26 / case 04

Analytics agents, router to picker

Text-to-SQL analysts and a root-cause agent for multi-brand retail, and what was left when the LLM router went.

Primary author, 478 of 515 commits on these paths

Python · LangGraph · DuckDB · SQL guardrails · FastAPI · Convex

root-cause turn
77.8 s→12 s
20,161 → 11,809 tokens, report unchangedmeasured
trust score
1 LLM call→0
made deterministic; a 12 s scan became sub-millisecondcommit
client identifiers in agent code
238→0
after moving config into knowledge packsdoc
tests
759→936
doc

My first year at Netision: conversational BI over a read-only warehouse. A sales analyst that plans, writes and guards SQL, injects row-level security and scores its own answer; a root-cause agent that explains why stores moved. In 2026 a two-level LLM router was put in front of nine specialist agents; four months later the product consolidated onto one chat backend, and the agents became tools inside its loop.

employer work: written up as architecture and decisions; client names, internal hosts and source stay private.

fig.how it fits together

01/04Context first: what was asked, which stores, which period. Only this step decides anything; everything after it is a writing task.

§ 1The motive

People in retail ask 'storewise sales versus last year' or 'why are these stores declining' in plain English, and expect a table they can trust, not a paragraph that sounds right.

§ 2Decisions

  • Deterministic code where a decision can be computed. The trust score became the share of the answer's bolded numbers that the fact sheet supports. One fewer model call per question.
  • Row-level security goes into the last WHERE. Not a subquery wrapper, because the store column usually isn't selected. Deny by default.
  • One codebase, many clients, no fork. A base pack plus client deltas, validated at boot. Collapsing the duplicates surfaced bugs nobody had seen, like an 'average' column that never matched and got summed.
  • From an LLM router to a picker. When the product moved to one chat backend where the user picks the agent, the classifier went: a dictionary is now the only place a name maps to a pipeline, and routing is deterministic as a side effect.

§ 3Ruled out

✗Batch every store family into one model call
Input fell 56%, but wall time went from 5.8 s to 15.8 s: one long serial decode. Kept the fan-out with a cacheable shared prefix.
✗Move to LangGraph to cut tokens
The waste wasn't orchestration. It was teaching the model the same rules twice.
✗Minimal reasoning effort everywhere
Low effort on the two prose stages took one step from 28.4 s to 8.2 s. 'Minimal' produced malformed output, and the deciding call was left alone.