Skip to content
← all work

Netision · 2025–26 / case 01

The query backend

The service that answers every question in an enterprise AI workspace: one capped agent loop, tools by contract, approval in the host, every token metered.

Primary author, about 87% of 1,620 commits

Python · FastAPI · asyncio · MCP · SQLite + sqlite-vec · DuckDB · Convex · Microsoft Graph · pytest

graded spreadsheet questions
2/8→6/8
after table discovery became progressive: index, then columns, then SQLeval
test modules
67
1,371 tests after the October consolidationcount
ledger events with usage and a price
246/246
live check after agents became tools (about 15 turns)measured

A FastAPI service behind the chat of an enterprise AI workspace. It resolves the documents a question is about, retrieves what it needs, runs a capped tool-calling loop over an MCP tool server, parks anything with side effects behind a human approval, and writes the answer as data the frontend renders. It started as a document drafter and became the core LLM backend of the product.

employer work: written up as architecture and decisions; client names, internal hosts and source stay private.

fig.how it fits together

01/07A question arrives with its model and its agent. Anything that can be refused is refused now, before the 202. A bad request fails as one, not as a broken answer later.

§ 1The motive

A workspace where people ask questions about their own documents, mail and data needs one place that decides what a question needs, calls the right tools, and can be trusted with the user's accounts. It also needs to say what each answer cost.

It began in May 2026 as a skill-based document drafter. From July I rebuilt it on a shared library I wrote (see the LLM library) into the backend that runs every chat turn, and merged in the analytics agents I had been building since July 2025.

§ 2How a turn runs

  • Validate before the 202. Unknown model, unknown agent, empty prompt: refused synchronously. A 202 followed by an error block reads to the user as the agent failing, not as a bad request.
  • Resolve and retrieve. History, tagged files and workspace documents. One embedding, top 12 chunks from this conversation's scope.
  • The loop. At most four rounds. Tool calls in a round run concurrently; every result is cut to 8,000 characters at one choke point. Then one call with no tools at all, so the model has to answer in words.
  • Persist. Invented links are stripped before the answer is saved, citations become chips, files are filed by id (never by URL), and the answer is one block write.
  • Meter. Every chat call, embedding, vision page, OCR page and generated image lands in a per-user ledger.

§ 3Decisions

  • Tags, not tool names. Tools declare produces_file, requires_approval, connector:<id>, format:<word> in their MCP metadata. New tools ship with one allowlist word, or no change here at all.
  • Credentials ride headers, never arguments. Tool arguments are model-controlled and re-sent in history. A model that could set a token could read another person's account.
  • Approval lives in the host. A send-mail call is parked on a card whose props are the call; the Send click runs it with a fresh token and no model in the path. I wrote the design; a teammate built the first version.
  • A skill is an extra prompt and nothing else. One system message after the history. Same loop, same tools. So a skill can't quietly change how files get made.
  • Agents are tools inside the loop. The analytics agents used to bypass the chat loop, so 'show me the sales' then 'mail it' re-ran the SQL. Now one loop holds connectors, files and agents.

§ 4Ruled out

✗Retry a call the guardrail blocked
A blocked call is billed its whole prompt, and a second one clears nothing.
✗Offer every file tool on every turn
Research and delivery compete for the same round budget. Gated on the words in the request, a news question stopped returning an unasked-for PDF.
✗Keep the warehouse on blob storage
Shipped, measured fast when warm, then reverted: it ran out of memory on a small VM with nowhere to spill.

§ 5What this doesn't fix

Conversation memory is process-local SQLite: fine at one or two queries a second per process, wrong for many replicas. That limit is written down where the next person will look, along with what would have to change.