Netision · 2025–26 / case 01
The query backend
The service that answers every question in an enterprise AI workspace: one capped agent loop, tools by contract, approval in the host, every token metered.
Primary author, about 87% of 1,620 commits
Python · FastAPI · asyncio · MCP · SQLite + sqlite-vec · DuckDB · Convex · Microsoft Graph · pytest
- graded spreadsheet questions
- 2/8→6/8
- after table discovery became progressive: index, then columns, then SQLeval
- test modules
- 67
- 1,371 tests after the October consolidationcount
- ledger events with usage and a price
- 246/246
- live check after agents became tools (about 15 turns)measured
A FastAPI service behind the chat of an enterprise AI workspace. It resolves the documents a question is about, retrieves what it needs, runs a capped tool-calling loop over an MCP tool server, parks anything with side effects behind a human approval, and writes the answer as data the frontend renders. It started as a document drafter and became the core LLM backend of the product.
employer work: written up as architecture and decisions; client names, internal hosts and source stay private.
01/07A question arrives with its model and its agent. Anything that can be refused is refused now, before the 202. A bad request fails as one, not as a broken answer later.
§ 1The motive
A workspace where people ask questions about their own documents, mail and data needs one place that decides what a question needs, calls the right tools, and can be trusted with the user's accounts. It also needs to say what each answer cost.
It began in May 2026 as a skill-based document drafter. From July I rebuilt it on a shared library I wrote (see the LLM library) into the backend that runs every chat turn, and merged in the analytics agents I had been building since July 2025.
§ 2How a turn runs
- Validate before the 202. Unknown model, unknown agent, empty prompt: refused synchronously. A 202 followed by an error block reads to the user as the agent failing, not as a bad request.
- Resolve and retrieve. History, tagged files and workspace documents. One embedding, top 12 chunks from this conversation's scope.
- The loop. At most four rounds. Tool calls in a round run concurrently; every result is cut to 8,000 characters at one choke point. Then one call with no tools at all, so the model has to answer in words.
- Persist. Invented links are stripped before the answer is saved, citations become chips, files are filed by id (never by URL), and the answer is one block write.
- Meter. Every chat call, embedding, vision page, OCR page and generated image lands in a per-user ledger.
§ 3Decisions
- Tags, not tool names. Tools declare
produces_file,requires_approval,connector:<id>,format:<word>in their MCP metadata. New tools ship with one allowlist word, or no change here at all. - Credentials ride headers, never arguments. Tool arguments are model-controlled and re-sent in history. A model that could set a token could read another person's account.
- Approval lives in the host. A send-mail call is parked on a card whose props are the call; the Send click runs it with a fresh token and no model in the path. I wrote the design; a teammate built the first version.
- A skill is an extra prompt and nothing else. One system message after the history. Same loop, same tools. So a skill can't quietly change how files get made.
- Agents are tools inside the loop. The analytics agents used to bypass the chat loop, so 'show me the sales' then 'mail it' re-ran the SQL. Now one loop holds connectors, files and agents.
§ 4Ruled out
- ✗Retry a call the guardrail blocked
- A blocked call is billed its whole prompt, and a second one clears nothing.
- ✗Offer every file tool on every turn
- Research and delivery compete for the same round budget. Gated on the words in the request, a news question stopped returning an unasked-for PDF.
- ✗Keep the warehouse on blob storage
- Shipped, measured fast when warm, then reverted: it ran out of memory on a small VM with nowhere to spill.
§ 5What this doesn't fix
Conversation memory is process-local SQLite: fine at one or two queries a second per process, wrong for many replicas. That limit is written down where the next person will look, along with what would have to change.