· 3 min read
You probably don't need multi-agent
The orchestrator, the critic, the researcher and the planner walk into your architecture. One model with good tools was already doing the job. The math and the exceptions.
#agents#architecture#cost
6 articles
· 3 min read
The orchestrator, the critic, the researcher and the planner walk into your architecture. One model with good tools was already doing the job. The math and the exceptions.
#agents#architecture#cost
· 2 min read
Frontier output pricing is down roughly 88% since early 2023. The clever cost engineering from two years ago is now complexity you pay to keep.
#cost#architecture#pricing
· 3 min read
Frontier models for frontier problems. For the other 80 percent of your pipeline, a small model at a fifth of the price is not a compromise, it is engineering.
#models#cost#architecture
· 3 min read
OpenTelemetry's GenAI conventions finally standardize token and cost telemetry. Instrument now or meet your bill the way we all met AWS bills, screaming.
#observability#cost#tooling
· 2 min read
Cache reads cost a tenth of normal input tokens. The math on when caching pays, when it doesn't, and the mistake that quietly doubles your write costs.
#cost#caching#api
· 4 min read
The decision tree for prompting, RAG and fine-tuning, in the order your wallet prefers. Most teams asking about fine-tuning need a better prompt and a dataset they don't have.
#fine-tuning#models#cost
get new writeups by email
no schedule, no spam. the next writeup, when it exists.