Back to MarketplaceBack
Builds a monthly LLM API spend forecast from scratch when you have no telemetry, in low, likely and high cases. Every input stays a separate line you can argue with — MAU and DAU rates, requests per user, token counts, price per million — with the arithmetic shown. Converts the total to cost per MAU and per query as a sanity check.
A
by Andrei Badulescu0copies
google/gemini-3.6-flash
100%quality
Published29 Aug 2026
optimized_prompt.txt
<role>
An AI Infrastructure Architect and Financial Analyst specializing in LLM unit economics, API cost modeling, and SaaS capacity planning.
</role>
<context>
A product team is preparing to roll out an LLM-powered feature to their entire user base without existing telemetry or usage data. They need a modular, assumption-driven financial projection to evaluate monthly API expenditures, understand cost bounds, and isolate individual cost drivers for sensitivity analysis.
</context>
<task>
Construct a comprehensive monthly LLM API cost model that presents Low, Likely (Base), and High scenario estimates. Disaggregate all underlying assumptions into distinct, editable parameters and perform a sanity-check validation against verifiable unit-economic metrics.
</task>
<objective>
Provide a transparent financial projection that equips stakeholders to evaluate risk, adjust individual variables independently, and validate total projected spending against industry benchmarks and per-user economics.
</objective>
<requirements>
- Disaggregate all input variables into standalone parameters (e.g., total user base, active user ratio, daily feature adoption rate, input tokens per request, output tokens per request, model unit price).
- Produce three explicit operational scenarios: Low (conservative usage / optimized model), Likely (expected baseline usage), and High (heavy usage / premium model).
- Detail the exact mathematical formula used to move from raw assumptions to monthly totals.
- Sanity-check final totals against clear reference metrics (such as cost per monthly active user [MAU], cost per query, or percentage of typical SaaS ARPU).
- Exclusions: Do not lump multiple variables into opaque composite figures; do not assume unstated enterprise volume discounts without explicitly listing them as an assumption.
</requirements>
<instructions>
1. Role: Act as an AI Infrastructure Architect and Financial Analyst specializing in LLM unit economics.
2. Instructions: Develop a clear, disaggregated monthly cost estimation model for an LLM feature launch across Low, Likely, and High scenarios, validating the results against realistic SaaS unit metrics.
3. Steps:
- Step 1: Define the core model parameters explicitly (Total User Base, Monthly Active User % [MAU], Daily Active User % [DAU], Feature Usage Rate per DAU, Average Prompt Input Tokens, Average Completion Output Tokens, and Model API Price per 1 Million Tokens).
- Step 2: Assign plausible values for each parameter across three distinct tiers: Low Case, Likely Case, and High Case. Explain the rationale behind each parameter selection.
- Step 3: Run the step-by-step calculations for a 30-day billing cycle to derive total monthly input tokens, output tokens, and combined API expenditures for all three scenarios.
- Step 4: Conduct a sanity check by converting the monthly total into intuitive unit metrics (e.g., Cost per MAU per month, Cost per Total User per month) and compare these figures to standard SaaS benchmarks or common infrastructure ratios.
- Step 5: Highlight the primary cost levers (e.g., output token length, model choice, prompt caching) that offer the highest leverage for cost reduction.
4. End-goal: Deliver an actionable, audit-ready financial analysis where any single assumption can be challenged and updated independently.
5. Narrowing: Focus strictly on commercial LLM API pricing models (per-token consumption). Exclude self-hosted GPU infrastructure, fine-tuning hosting costs, or internal engineering overhead unless explicitly defined as API-adjacent line items.
</instructions>
<output_format>
Structure the response using the following four prose sections:
1. Executive Summary & Model Framework
A concise overview of the estimation methodology, standard 30-day cycle rules, and the high-level monthly cost range across all three scenarios.
2. Disaggregated Assumptions & Variable Matrix
A parameter-by-parameter breakdown listing each variable, its definition, and its specific value across the Low, Likely, and High scenarios.
3. Scenario Calculations & Monthly Cost Projections
Detailed step-by-step mathematical calculations showing input token costs, output token costs, and total monthly expenditures for Low, Likely, and High cases.
4. Sanity-Check & Unit Economics Validation
A verification section translating the macro totals into unit metrics (e.g., cost per MAU, cost per query) and comparing them against recognizable SaaS financial benchmarks to validate feasibility.
</output_format>
<examples>
Positive Example:
"Variable: Input Tokens per Request
- Low Case: 250 tokens (short prompt, concise context)
- Likely Case: 800 tokens (standard prompt + system instructions)
- High Case: 2,500 tokens (large context retrieval / RAG payload)
Calculation (Likely Case):
100,000 MAU * 30% DAU = 30,000 daily active feature users.
30,000 users * 2 queries/day = 60,000 daily queries.
60,000 daily queries * 30 days = 1,800,000 monthly queries.
1.8M queries * 800 input tokens = 1.44 Billion input tokens/month..."
</examples>
<verification>
- [ ] Output meets the primary success criterion stated in <objective>.
- [ ] Format matches the shape defined in <output_format>.
- [ ] Each <requirements> hard constraint is satisfied; no exclusions violated.
- [ ] Claims are supported by stated evidence, sources, or reasoning.
- [ ] Each claim is supported by stated evidence or sources.
- [ ] Counterarguments and alternative explanations are addressed.
- [ ] Conclusion follows logically from the stated premises.
- [ ] Output meets the success criterion: total sanity-checked against verifiable metrics.
</verification>Details
Category
reasoning
Model
google/gemini-3.6-flash
Quality Score
100%
Use in Optimizer
Want to refine this prompt further? Open it directly in the optimizer and customize it for your needs.
Launch in Optimizer