Netision · 2026 / case 02
One library, not four copies
The shared LLM layer every backend installs: 17 models behind one interface, a Stop that actually stops, and a billing-grade ledger.
Created and sole maintainer, 100 of 100 commits
Python 3.12 · uv · asyncio · provider APIs · DuckDB · sqlite-vec · ClickHouse · OpenTelemetry
- releases
- 86
- tags on the shared remote, v0.1.0 → v0.59.1, July–October 2026count
- chat models behind one interface
- 17
- 6 providers in use, 8 provider routes, plus 2 image providerscount
- guardrail blocks on the same payload
- 3/12→0/12
- one confidentiality line instead of threemeasured
- pricing rewrite, golden cases identical
- 308/308
- an equivalence check against cases captured before the changecommit
The model config had been copy-pasted into four repositories and had drifted. One copy was about 2,500 lines away from another. I pulled it into a versioned Python library: published by git tag, never deployed. It now holds the model switchboard, document ingestion, the table store, cancellation, observability and the cost ledger. 86 releases in three months.
employer work: written up as architecture and decisions; client names, internal hosts and source stay private.
01/05The frontend sends a model key with the question. It is set once, at the request boundary. So no function in between grows a model argument it only forwards.
try it
Where does an upload go?
Pick a file and watch the document router decide: by what's in it, not what it's called.
upload · “Upload: a pitch deck, all pictures and almost no text.”⎘ board-review.pptx
t = 0 ms · timings illustrative- accepted.pptx · under 20 MB
- tabularno
- density52 chars, whole deck
- vision6 pages in flight
- chunksmarkdown · 1200/200
starting…
§ 1The motive
A fix or a new model added in one repository never reached the others. This is a library, not a service: published as a git tag, installed at build time, configured by the app at runtime. Apps pin an exact version, so a bad release can never cascade until an app opts in.
§ 2Decisions
- Model chosen per request, through a context variable. Set once at the request boundary, so nothing in between threads a model argument it doesn't use.
- Per-model settings live beside each model. One environment variable used to tune every model at once, which is how one model got mistuned to satisfy another.
- No house output cap. A global 4,096-token cap cut long structured answers mid-object; a larger one made every Claude call fail. Each model now gets its own maximum.
- Stop is a
BaseException. Pipelines are full ofexcept Exceptionblocks that turn failures into error messages. And a stop is not an error. The proxy checks before every model call with no wiring at the call site; because a cached response skips that check, pipelines also check again right before they publish. - Route documents by content, not extension. Below 150 characters a page, a file is pictures: it goes to vision transcription once, at upload. Two real decks had 52 and 168 characters of text in total.
- Free, unpriced and unmeasured are three different states. A zero for a payload nobody could read is how spend goes missing quietly.
§ 3Ruled out
- ✗Comment out models that stopped answering
- A dead entry reads as a working option and costs the caller a 404 at runtime, not at deploy. Seven of eight entries were dead when I checked; now every entry is called live before it's listed.
- ✗Filter vector search by scope after the search
- A two-chunk file in a 435-chunk store returned nothing. Scope became a partition key; existing stores migrated without re-embedding anything.
- ✗Trust the model's SQL because the prompt says 'read only'
- Two independent defences instead: a locked-down connection and a statement check. Either alone would probably hold. Neither alone is worth betting a filesystem on.
§ 4What this doesn't fix
Every change costs a release, so apps get register_model() and reader registries for things only they need. Some models are wired but unverified in production, and they're labelled that way.