I measured my semantic cache. Then I turned it off.
A semantic cache saves money by answering a repeat question from memory instead of calling the model again. Mine looked excellent against paraphrases and fell apart against near misses. Here is the measurement, the number that ended it, and what I would build instead.
1 update· latest August 26, 2026