note · 9 October 2026 · 5 min readdraft
Spreadsheets are not documents
Why chunk-and-embed fails on spreadsheets, what replaced it, and the two independent locks on SQL a model writes.
The first version treated a spreadsheet like a PDF: chunk it, embed the chunks, retrieve the nearest ones. It was very good at “what does row 400 say”. Nobody asks that.
People ask “which region grew fastest”, “total spend by channel”, “compare this plan with last quarter's”. Those are sums, group-bys and joins. Similarity search doesn't add up columns.
A sheet is not a table
Real files don't look like CSVs. A media plan is a sheet containing tables: side by side, stacked under banner rows, with two-level headers and totals in the middle. Measured on four real plans, a standard reader returned one column and zero rows for a side-by-side sheet; another returned twenty-three columns called Unnamed. Neither is a bug in those libraries. They answer “read the table”, and the question here is “find the tables”.
So detection comes first: split on blank rows and spacer columns, treat banner rows as titles, fold two-level headers into names, and flag total rows by their label and by arithmetic, because not every total says it's one.
Query, don't retrieve
Each detected table lands in DuckDB, scoped to the conversation, and only a short schema card per table is embedded. The model discovers progressively: the index, then columns on request, then SQL. An ordinary chat never pays for any of it. The tool is only offered when a data file is attached.
On a graded set of eight real questions: 2/8 before, 6/8 after. The jump came once discovery stopped overrunning the 8,000-character cut on tool results.
Totals read double
One plan summed to exactly twice its real figure: the total rows were being added to the rows they total. Because totals are flagged at upload, the query excludes them. The sum is corrected in the query, not in the answer.
A related trap: columns that look alike but aren't. Potential, planned and achieved reach must never be added together, and a rate is weighted by what it's a rate of. An average of averages is silently plausible and wrong. So the column cards carry a concept, a qualifier and an aggregation rule, decided by code, not by the model.
The SQL is untrusted
DuckDB can read files. read_csv('/etc/passwd') is valid SQL. Any SQL a language model writes is untrusted input, no matter how the prompt is worded. So there are two independent defences:
- the connection is locked down before it's used (no external access, configuration frozen), and the store refuses to open if it can't verify that;
- every statement is checked: exactly one, read-only, no file functions, and only tables that belong to this conversation. The catalog table is excluded too. One
SELECT * FROM _catalogwould leak every other conversation's schema.
Either alone would probably hold. Neither alone is worth betting a filesystem on.
The refusals are written for the model, not for a log: they say what was wrong and what's allowed, so the next round fixes the query instead of giving up.
written from my own design docs, commit messages and tests; the numbers are theirs.
next note →Never hand the model the keys