Skip to content

thinking

How I think

Rules I work by, each with the incident that produced it, and the shape I keep reaching for when I start a repository.

§ 01how I think

Rules I work by

Each one is a rule plus the thing that happened to produce it. I keep them in a decision log next to the code. These are the ones I'd defend at a whiteboard.

  1. what happened

    Chat spend was being logged against a user named after the service itself. Embeddings, OCR pages and generated images weren't metered at all, and one model was costed at a fifth of its invoice.

    what I do now

    Numbers come only from the provider's own usage payload. A token is never estimated. A payload nobody could read is flagged usage_missing; a model with no rate is flagged pricing_missing; unpriced models are named at boot, before the first invoice.

    ∴ Free, unpriced and unmeasured are three states, and they look different on every report.from one library, not four copies →

  2. what happened

    Tool arguments are written by the model and re-sent in history. Anything that rides in them (a user id, a token, an API host) is something the model could set.

    what I do now

    Identity, credentials, hosts, approvals and SQL scope are enforced by code and headers, not by sentences in a prompt. A model that could set its own model could pick anyone's.

    ∴ No tool takes a user, a token or a host, and a test fails if one ever does.from the mcp tool server →

  3. what happened

    An async API that answers 202 and then fails in the background looks, to the person using it, like the agent breaking, not like a bad request.

    what I do now

    Unknown model, unknown agent, empty prompt: refused synchronously, before anything is promised. Inside the turn, the opposite: a tool server that's down means answering without tools, not dying.

    ∴ Frontend renames that used to drop fields silently now fail loudly, on the first request.from the query backend →

  4. what happened

    The model writes SQL over uploaded spreadsheets, and the engine can read the filesystem: read_csv('/etc/passwd') is valid SQL.

    what I do now

    A connection that is locked down before it's used, and a statement check: one read-only statement, only this conversation's tables, no file functions. Refusals are worded for the model, so it fixes the query next round.

    ∴ Either alone would probably hold. Neither alone is worth betting a filesystem on.from one library, not four copies →

  5. what happened

    Pressing Stop logged a cancel, and the stopped query still wrote its result.

    what I do now

    Cancellation is a BaseException, so it sails past the except Exception blocks that turn failures into error messages. The model proxy checks before every call. But a cached response skips that check, so each pipeline also checks right before anything reaches the user.

    ∴ Three rules, pinned by their own test file.from the query backend →

  6. what happened

    Batching every store family into one model call looked like an obvious saving: 56% less input. Measured, that step went from 5.8 s to 15.8 s: one long serial decode instead of four running at once.

    what I do now

    Every lever gets measured on the real thing before it's trusted, and a rejected option is written down with its numbers next to the decision, so nobody re-runs the experiment by accident.

    ∴ The fan-out stayed. The turn got fast elsewhere (slow views stored as tables, the date range read once, lower effort on the two prose stages): 77.8 s → 12 s.from analytics agents, router to picker →

  7. what happened

    Every round of an agent loop re-sends every earlier round's results. A 500 KB PDF returned as base64 would cost about 170,000 tokens per turn.

    what I do now

    A hard cap on rounds. Results truncated at one choke point, not per tool. Files by reference (~30 tokens). Tools offered only when this request can use them. Prefixes kept stable so the cache can do its job.

    ∴ Later rounds re-read the prefix from cache at about a tenth of the price.from the query backend →

  8. what happened

    0 of 35 registered sources were cited in one logged turn. The instruction to cite was gated differently from the registration. And where it was sent, it sat above seven searches' worth of results.

    what I do now

    The model writes its answer from the bottom of the message list up. So notes go directly under the results they're about, and standing instructions go after the history, because above it they lose.

    ∴ One predicate now decides both whether sources are registered and whether the note is sent, and the note sits directly under the results.from the query backend →

  9. what happened

    The model config lived in four repositories. One copy was about 2,500 lines away from another; a fix in one never reached the rest.

    what I do now

    One library, published by tag. Registries that answer the question instead of lists that drift. Contract tests that read the TypeScript frontend from the Python backend, so the two can't disagree quietly.

    ∴ 86 releases in three months; the two production services install it instead of keeping a copy.from one library, not four copies →

  10. what happened

    A guardrail blocked some calls and not others, and only through the proxy. Fifty-five direct calls reproduced nothing. A payload assembled here is not the payload sent.

    what I do now

    Until the symptom has been flipped in both directions, it's a hypothesis, and it gets labelled one: confirmed, probable or hypothesis, never “should work”.

    ∴ Three copies of one line: 3/12 blocked. One copy: 0/12.from one library, not four copies →

  11. what happened

    Diagnostic tables had grown without limits. One column of payloads was 71% of the database, for 735 distinct values across 1,571 rows; another table had never been written.

    what I do now

    Measure first (anti-joins, blank rates, row-by-row comparisons), then drop. A column that reads 0 everywhere says “none happened”, when the truth was “not counted”. A module nobody could call was speculation, so it goes.

    ∴ One rule for retention anyone can state, and one owner for every fact.from one library, not four copies →

  12. what happened

    Two halves of a document reader each passed their tests for months, while a whole class of decks was refused. Each half was right alone.

    what I do now

    Tests are named as behaviours: a stop during the final call does not publish, an unpriced model is flagged, not free. Paired failures get paired tests. Snapshots get replaced by properties: a recording only says “unchanged since someone recorded it”.

    ∴ The suite reads like the rules it enforces.from the query backend →

§ 02how I structure work

The shape I keep reaching for

Not a framework. Habits. The same five or six decisions show up in every repository I start, because each one was learned the expensive way first.

  • repo/
  • CLAUDE.mda decision log: each rule, then the incident behind it
  • docs/plans/the plan comes first, then one commit per step
  • src/
  • api/handlers orchestrate, and do nothing else
  • core/registries, not branches
  • tests/named as behaviours: “a stop during the final call does not publish”
  • .claude/skills/checklists for the coding agents I work with
  1. P-01

    Plan first, then one commit per step.

    A plan document written against the code, with file and line references, before anything changes. Big plans get adversarial reviews first. The connector plan had six. Then the plan is executed one step per commit, so the history reads like the plan.

    ∴ working with Claude Code, the observability control plane went through eight planned phases in one day

  2. P-02

    A decision log next to the code.

    My project instructions aren't a style guide. Each paragraph is a rule in bold, then the incident or measurement that produced it, so the next person (or agent) knows which rules are load-bearing. Comments are one or two lines and say why, not what.

    ∴ about 700 lines of rules, each with its reason

  3. P-03

    Agents get checklists, not vibes.

    I write skills for the coding agents I work with: a release gate in tiers, a root-cause routine that labels findings confirmed, probable or hypothesis, a token audit, a dependency audit. Each one ends the same way: merging is never the agent's call.

    ∴ seven project skills in the query backend

  4. P-04

    Registries, not branches.

    Things that grow (models, tools, file readers, skills) register themselves, and the code asks the registry instead of keeping a list. Adding a model is a line of data; adding a tool is a folder.

    ∴ a file type was refused at the door while a working reader sat behind it, until the gate asked the registry