Delivered twice, to two audiences. For leadership it is a financial modelling session: a twelve-month projection, cloud versus on-premises TCO, and where the money actually goes. For engineers it is caching, routing, compression, batching — and an honest account of when not to use an LLM at all.
What we cover
- Why LLM costs spiral: token economics & agentic loops
- Building your own 12-month cost projection model
- Cloud vs. on-prem: GPU capex, throughput, TCO analysis
- Right-sizing model selection by task complexity
- Prompt compression: system prompt caching, output control
- Caching strategies: semantic, exact-match, TTL policies
- Batching, async processing & request coalescing at scale
- When NOT to use an LLM: rules, classifiers, retrieval-only
What participants leave able to do
- Build a rigorous cost model before committing infrastructure budget
- Measure cost and quality trade-offs before changing models or architecture
- Eliminate redundant LLM calls with semantic & exact-match caching
- Evaluate cloud vs. on-prem with a structured financial model
- Design agentic systems with built-in cost guardrails
- Replace costly LLM calls with cheaper non-generative alternatives
- For
- Engineering leads · AI architects · CTOs · FinOps teams
- Format
- Half day (3 hrs)
- Track
- Practitioner Upskilling · Executive & Leadership