· 2 min read
The price crash quietly turned your 2024 cost hacks into debt
Frontier output pricing is down roughly 88% since early 2023. The clever cost engineering from two years ago is now complexity you pay to keep.
#cost#architecture#pricing
6 articles
· 2 min read
Frontier output pricing is down roughly 88% since early 2023. The clever cost engineering from two years ago is now complexity you pay to keep.
#cost#architecture#pricing
· 3 min read
Frontier models for frontier problems. For the other 80 percent of your pipeline, a small model at a fifth of the price is not a compromise, it is engineering.
#models#cost#architecture
· 2 min read
Cache reads cost a tenth of normal input tokens. The math on when caching pays, when it doesn't, and the mistake that quietly doubles your write costs.
#cost#caching#api
· 1 min read
Schema-enforced output killed the regex parser. It did not kill the need to version, validate and monitor what your model returns.
#structured-outputs#api#reliability
· 4 min read
Million-token windows did not fix context rot, they subsidized it. How I split the window into line items and stopped my agents from getting dumber mid-task.
#context#agents#architecture
· 4 min read
The decision tree for prompting, RAG and fine-tuning, in the order your wallet prefers. Most teams asking about fine-tuning need a better prompt and a dataset they don't have.
#fine-tuning#models#cost
get new writeups by email
no schedule, no spam. the next writeup, when it exists.