slowbench

Measure first. Then decide.

Most of what I believed about my own systems did not survive measurement.

Notes on retrieval, evaluation, and the systems around them — measured before written. Every number here comes from a run I did myself, against the system that actually serves traffic — not a copy made for the benchmark.

Retrieval

0

Code search, embeddings, and ranking.

Inference

0

Local inference, Metal/GPU, and batch tuning.

Evals

0

Measurement methodology and benchmark design.

Agents

0

Coding agents and context management.

Recent

Nothing here yet.