Scenario Notes — travels with shared links
Optimized Stack · Based on Your Budget
At $200/month, your best play is Gemini 2.5 Flash + light RAG.
A single budget-tier API model with prompt caching gives you the most quality per dollar at this constraint. Self-hosting doesn't earn its keep until you cross $800/mo or have GPUs already in-house.
Monthly Allocation
What You Can Run
Token Estimate · Heuristic
Per-call cost across every model.
A character-based heuristic (~4 chars/token, denser for CJK). Expect ±15-20% vs. the real provider tokenizer.
System + Tools
0
tokens
User + History
0
tokens
Output
500
tokens
Total / Call
500
in + out
At 1,000 calls/day, this is $0/mo on the cheapest model and $0/mo on the most expensive.
§ I · Per-Model Cost Breakdown
| Model | Tier | Speed | $ / Call | $ / Day | $ / Month |
|---|
Self-hosted rows amortize the monthly GPU baseline across your stated call volume. API rows reflect cache discount on input tokens.
§ II · Ready-to-Run Snippets
cURL
Python SDK
Total Daily Tokens
10M
tokens / 24h
Monthly Token Volume
300M
tokens / 30 days
Annual Token Volume
3.65B
tokens / year
§ I · Per-Model Cost Comparison
Small Model · Self-Hosted
Phi-4 Mini · 3.8B
Monthly TCO
$500
$6K / year · $0.50 per 1M tokens
Daily Cost$15
Context Window8K
Inference SpeedVery Fast
Eng. Complexity
Cross-Module IQ
CRUD operations, FAQs, deterministic flows — the disciplined laborer.
Medium Model · Self-Hosted
DeepSeek V3.2 · 14B
Monthly TCO
$2,000
$24K / year · $4 per 1M tokens
Daily Cost$70
Context Window32K
Inference SpeedFast
Eng. Complexity
Cross-Module IQ
Module copilots, workflow automation — the balanced workhorse.
Frontier Model · API
Claude Sonnet 4.6
Monthly TCO
$6,000
$72K / year · $17.50 per 1M tokens
Daily Cost$200
Context Window200K
Inference SpeedMedium
Eng. Complexity
Cross-Module IQ
Enterprise orchestration, analytics, deep reasoning — the strategist.
§ II · Cost Curves & Hybrid Allocation
Monthly cost, by model class
USD / month — at current volume
Hybrid workload split
Optimal token allocation across tiers
$2.1K
Hybrid / mo
Small 60%
CRUD, FAQs, deterministic flows
Medium 30%
Module copilots, summarization
Large 10%
Cross-module reasoning, planning
§ III · Strategic Verdict
Recommendation Engine · Live
Run a hybrid stack — and you save ~$3.9K/mo over pure frontier.
At 100 users × 100K tokens/day, a pure frontier-model strategy costs roughly $6K/month. A hybrid allocation — Phi-3 for the boring 60%, DeepSeek 14B for the middle 30%, GPT-4o/Claude for the strategic 10% — delivers the same end-user experience for ~$2.1K/month. The catch: hybrid demands an orchestration layer (MCP routing + intent classification) which adds roughly 0.4 engineer-years upfront. Payback under 4 months at this volume.
§ IV · Hidden Costs & Insights
$48K
Hardcoding penalty (yr)
Small models need more workflow logic. At current engineer rates, that's the implicit annual tax on choosing self-hosted small alone.
14 mo
Break-even: small vs large
Time it takes self-hosted savings to pay back the extra engineering effort vs going pure-API.
$47K
Annual savings (hybrid)
What you keep by routing intelligently vs sending every token to a frontier API.
$0
Prompt cache savings / mo
Cached input tokens cost 90% less. At current cache hit rate, this is what you save vs zero caching on the frontier model.
$0
RAG infra / mo
Vector DB + embedding generation. The line item most calculators miss — added to every strategy uniformly.
$0
Peak load surcharge / yr
Self-hosted GPUs must size for peak, not average. This is the extra infra cost above a flat-load baseline.
§ V · Full Comparison Matrix
| Metric | Phi-3 Mini | DeepSeek 14B | GPT-4o / Claude | Hybrid |
|---|