Frontier models for frontier problems. For the other 80 percent of your pipeline, a small model at a fifth of the price is not a compromise, it is engineering.
OpenTelemetry's GenAI conventions finally standardize token and cost telemetry. Instrument now or meet your bill the way we all met AWS bills, screaming.
Cache reads cost a tenth of normal input tokens. The math on when caching pays, when it doesn't, and the mistake that quietly doubles your write costs.
The decision tree for prompting, RAG and fine-tuning, in the order your wallet prefers. Most teams asking about fine-tuning need a better prompt and a dataset they don't have.
#fine-tuning#models#cost
~$ a few cookies for readership stats. that ok?
~$ get new writeups by email
no schedule, no spam. the next writeup, when it exists.