Measure first. Then decide.
Most of what I believed about my own systems did not survive measurement.
Notes on retrieval, evaluation, and the systems around them — measured before written. Every number here comes from a run I did myself, against the system that actually serves traffic — not a copy made for the benchmark.
Retrieval
0Code search, embeddings, and ranking.
Inference
0Local inference, Metal/GPU, and batch tuning.
Evals
0Measurement methodology and benchmark design.
Agents
0Coding agents and context management.
Recent
Nothing here yet.