- Token price per million is not the same as feature cost per customer—infrastructure, caching, vector databases, and embeddings often account for 40–60% of total spend and never appear on your API invoice.
- Nearly 8 in 10 IT leaders report unexpected charges from consumption-based AI pricing; tracking average tokens per user per month by plan is the only way to prevent margin collapse.
- Platform vendors lock you in by changing pricing models mid-contract, and rebuilding to switch takes longer than losing customers; the rebuild-vs-buy decision depends on consumption stability and your negotiating power.
Your AI feature is hemorrhaging money because you're optimizing for the wrong metric. You found the cheapest model at $0.03 per million input tokens and $0.13 per million output tokens, but you have no visibility into whether that feature is actually profitable. That's the gap between per-token pricing and per-customer margin—and it's where almost every founder gets surprised.
The problem starts with what the invoice doesn't tell you. When you call an LLM API, you pay for tokens. When you build a real feature, you pay for tokens plus everything else: vector databases, embedding generation, reranker calls, semantic caching, orchestration runtime, fallback logic, and error handling. In production RAG and agentic deployments, infrastructure surrounding the model call routinely represents 40 to 60% of total feature spend, and almost none of it shows up as a line item from your model vendor.
That gap matters because it's where your unit economics break. For SaaS companies, allocating AI token cost per customer is essential for understanding margin—if AI costs per customer exceed what you're charging, you have a pricing problem that won't fix itself. But you can't fix what you can't measure.
The Real Cost Breakdown
Simple cost models suggest that simple LLM features cost $5K–$20K, RAG systems $20K–$60K, AI agents $25K–$80K, and legacy integrations $40K–$150K+ to build. Those numbers cover initial development. Ongoing costs are where founders get blindsided.
Token consumption varies wildly by use case. A question-answering bot might burn 500–1,000 tokens per user per month. A content generation feature for a creator platform might burn 50,000. A data analysis agent running daily jobs might burn 500,000. Your per-token cost is irrelevant if you don't know what "per customer" actually means for your specific product.
Beyond tokens, infrastructure costs scale independently. A vector database holding 100,000 embeddings costs roughly $200–500 per month. A semantic caching layer adds $100–300 per month. A reranker for RAG adds another $50–150 per month. These are fixed or semi-fixed costs that don't scale linearly with token volume. If you have 50 customers, you're dividing that by 50. If you have 5,000, it's cheaper per customer—but the bill doesn't shrink.
Model drift costs $3,000–$10,000 per month ongoing, scaling with data volume and model complexity. That's the cost of monitoring whether your model's outputs are still accurate, retraining when they drift, and paying for that retraining. It's not optional if you care about quality.
| Cost Category | Per-Month Range | Visibility |
|---|---|---|
| Token consumption (model API) | $500–$50K+ | On invoice |
| Vector database hosting | $200–$500 | Separate vendor |
| Embedding generation (if external) | $100–$300 | Separate vendor |
| Caching & orchestration | $100–$400 | Server logs, hard to isolate |
| Model drift & monitoring | $3K–$10K | Internal labor or vendor |
| Fallback logic & redundancy | $200–$800 | Infrastructure costs |
Why Token Opacity Breaks Your Unit Economics
You can't negotiate your API contract intelligently if you don't know how many tokens you're actually consuming. You'll sign a $50K monthly cap thinking you're safe, then hit it three months later when a customer runs an agent for the first time. You can't forecast because you don't have a baseline of tokens per customer or per feature. You can't optimize usage because you don't know which features are burning tokens wastefully.
The industry knows this. In 2025, organizations spent an average of $1.2 million on AI-native applications, and nearly 8 in 10 IT leaders report being hit with unexpected charges tied to consumption-based AI pricing. That's not a bug in a specific product. That's a systemic signal that token-based pricing is harder to manage than anyone selling it admits.
Consumption Tracking: What You Actually Need
Start with three numbers:
- Average tokens per customer per month, by plan tier. If your Pro plan customers use 10,000 tokens per month on average and your Enterprise customers use 500,000, your margin structure is completely different for each. You need to know this before you price anything.
- The 90th percentile user. What does your top 10% of users look like? Are they using the AI feature 50 times as much as the median? If so, one power user could blow through your monthly budget. You need headroom for them or a way to throttle their consumption.
- Token cost per feature. If your chat feature costs $0.50 per user interaction and your content generation feature costs $3.00, you need to know that. You can't optimize a feature if you don't measure it.
Build this instrumentation into your product before your AI feature goes live. Every API call should log: customer ID, plan tier, feature name, model used, input tokens, output tokens, latency, and cost. Aggregate weekly and track trends.
The Vendor Lock-In Problem
This is the ugliest hidden cost. You optimize your feature for Claude, build workflows around its specific prompt behavior, and integrate its API deeply into your product. Then Anthropic changes their pricing from $0.003 to $0.01 per 1K input tokens. Your per-customer AI cost just went up 3x overnight. Switching to another model takes 4–8 weeks if you're fast. Your customers churn in the meantime.
Or OpenAI stops offering the per-token model you architected around and moves to per-seat pricing. You're locked in because you've optimized your code and your cost structure around their old terms.
This is where the rebuild-vs-buy question gets real. Rebuilding an AI feature to use a different vendor can take 4–8 weeks and cost $30K–$100K in engineering time. Accepting a 30% price increase might be cheaper than rebuilding. But if the vendor knows you're locked in, they can price accordingly.
Rebuild vs. Buy: The Real Decision
You have three options when your AI costs blow up:
1. Keep buying from the vendor. Accept the price increase, negotiate if you have leverage (you usually don't), and pass costs to customers if your contract allows. This works if the cost increase doesn't erase your margin.
2. Migrate to a cheaper vendor. This takes 2–6 weeks for a simple feature, 6–12 weeks for a RAG system. You'll need to retune prompts, test outputs, possibly retrain models. Cost: $20K–$80K in engineering time. Only do this if the annual savings exceed the migration cost.
3. Build your own inference layer. Host open-source models (Llama 2, Mistral, Qwen) on your own infrastructure. This removes vendor lock-in but increases operational complexity. You're now managing model serving, GPU costs, fine-tuning, and monitoring accuracy. Fixed monthly cost: $500–$5K depending on model size and traffic. Variable cost: negligible compared to API calls. Only makes sense if your AI token consumption is $20K+ per month and you have ops capacity to manage it.
The decision hinges on three factors:
- Monthly AI cost. Below $5K, migrate or accept the increase. $5K–$20K, evaluate self-hosting. Above $20K, self-hosting almost always makes economic sense if you can stomach the ops overhead.
- Token consumption stability. If your feature's token burn is predictable and flat month-to-month, you can forecast and absorb price changes. If it's volatile (agents running variable-length tasks, user-generated workloads), vendor increases hit harder.
- Negotiating power. If you're a $10M ARR company using Claude for a core feature and spending $100K per month, Anthropic will negotiate. If you're a $1M ARR company spending $2K, you take what they offer.
Practical Steps Today
Week 1: Audit all AI vendor contracts. Write down the pricing model, usage caps, and what happens if you exceed them. Note any clauses about price changes.
Week 2: Instrument your product to log tokens per customer per feature. If you're on Vercel, use their built-in analytics. If you're custom, add logging to every model API call.
Week 3: Calculate your average tokens per customer per month for each plan tier. Multiply by your vendor's per-token cost. Compare to what you're charging. If your AI cost is more than 20% of what you charge customers, you have a pricing problem.
Week 4: Model the cost of migrating to a cheaper vendor or self-hosting. Compare to your annual vendor cost. If migration saves $100K+ per year, start planning the transition.
The median closed-model launch cost fell from $6.00 in 2024 to $4.38 in 2025 and $3.75 in 2026, so the raw cost to call a model keeps dropping. But your total AI cost—tokens plus infrastructure plus drift plus lock-in—keeps rising because you can't see it.
FAQ
How do I know if my AI costs are out of control?
Calculate your AI cost per customer per month and compare it to what you charge them. If it's more than 15–20% of their subscription price, you have a margin problem. If you don't know this number, start tracking it immediately. You're flying blind.
Should I self-host open-source models or keep using APIs?
If your monthly AI token cost is consistently above $15K–$20K per month, self-hosting becomes economically sensible. You'll trade vendor lock-in risk for operational complexity. Only do this if you have infrastructure ops capacity or can hire it. If your token consumption is under $10K per month, APIs are cheaper because you don't pay for idle GPU capacity.
What if a vendor changes their pricing mid-contract?
Read your contract. Some allow unilateral price changes with notice; others don't. If they do, you have 30–90 days to decide whether to stay or migrate. Plan your migration contingency before it happens. Have a backup vendor ready. Negotiate volume discounts if you're large enough.
How much does it cost to migrate from one LLM API to another?
Simple features (Q&A, summarization): $5K–$15K, 2–4 weeks. RAG systems: $20K–$50K, 4–8 weeks. Complex agents: $40K–$100K+, 8–12 weeks. The cost depends on how deeply your product is integrated with the original vendor's API, prompt tuning required, and testing thoroughness needed.






