Frontier models for frontier problems. For the other 80 percent of your pipeline, a small model at a fifth of the price is not a compromise, it is engineering.
Cache reads cost a tenth of normal input tokens. The math on when caching pays, when it doesn't, and the mistake that quietly doubles your write costs.
Million-token windows did not fix context rot, they subsidized it. How I split the window into line items and stopped my agents from getting dumber mid-task.
The decision tree for prompting, RAG and fine-tuning, in the order your wallet prefers. Most teams asking about fine-tuning need a better prompt and a dataset they don't have.
#fine-tuning#models#cost
~$ a few cookies for readership stats. that ok?
~$ get new writeups by email
no schedule, no spam. the next writeup, when it exists.