Azure AI planning tool

Cost Model

Estimate token demand, pay-as-you-go spend, tiered model routing, and provisioned throughput.

Auto-saved locally
Live calculation

User assumptions

Adoption and working pattern

Core estimate

Token scenarios

Define the prompts, outputs, model, and share of transactions for each workload.

RAG document estimator

Add document context only when retrieval is part of the solution.

Optional

Monthly PAYG forecast

Model strategy

Single model or scenario-weighted routing

Forecast detail

Monthly demand, cost, and throughput

Scenario comparison

Each enabled workload shown independently at 100% of baseline transaction volume.

How forecast costs are calculated

PAYG: (input tokens × input price + output tokens × output price) ÷ 1,000,000.

Routing: each enabled scenario uses its transaction allocation and selected model. Model Router cost is input tokens × router input price ÷ 1,000,000.

The selected quarterly growth rate compounds every three months across the full 36-month forecast.

Peak weighting

Concurrency buffer applied to average TPM

PTU recommendation

Provisioning respects each model's minimum and incremental PTU rules

Cost comparison

Baseline month and 36-month totals

PAYG versus PTU break-even

Quarterly totals use monthly forecast demand and a fixed monthly PTU commitment.

QuarterPAYGPTUDifferenceLower cost
How capacity is calculated

Peak TPM: user-based demand uses average TPM × peak multiplier. Transaction-based demand distributes weekday and weekend transactions across four-hour windows and takes the highest resulting TPM.

Required PTU: peak TPM ÷ model TPM per PTU, rounded up. Provisioned PTU then rounds up to the model's minimum and increment.

PTU cost: provisioned PTU × monthly unit price. Reserved pricing uses the workbook's reserved monthly unit price.

Azure AI pricing

GBP per million tokens and PTU assumptions

Add the model or models you want to price manually using Add model, including PTU details where required. Currency changes do not convert existing values, so enter prices in the selected currency or use Refresh official data to replace available PAYG prices.

Refresh updates PAYG prices for the selected deployment scope from the Azure Retail Prices API. Data Zone and Regional prices are filtered by Azure region. Supported Global and Data Zone throughput assumptions come from the Microsoft Foundry PTU sizing table. PTU costs, reserved rates, Model Router, regional throughput, and unmatched models remain manual.

ModelInput GBP/MOutput GBP/MTPM/PTUMin PTUPTU incrementPTU monthly (GBP)Reserved monthly (GBP)

Indicative estimate only. Actual Azure costs may differ. Users should validate all assumptions, token volumes, pricing and calculated outputs against current Azure pricing and applicable commercial agreements before making financial or architectural decisions.

Estimate prompt tokens

1 token ≈ 0.75 words100 words ≈ 130 tokens1 page ≈ 500–700 tokens
Words 0Estimated tokens 0

Using the cost model

From workload assumptions to a reviewable cost forecast

Work through the tabs from Estimate to Capacity. Results update as you edit and are saved only in this browser.

  1. Choose the forecast basis

    Select User based for adoption-led demand or Transaction based for a known monthly workload. Enter the starting volume and quarterly growth as a percentage or factor.

  2. Describe each workload

    Add a scenario for every distinct prompt pattern. Enter system, user, history, tool, and expected output tokens, or use Add text to estimate tokens from representative content. Set enabled scenario allocations to total 100%.

  3. Add retrieval assumptions when needed

    Open the optional RAG document estimator only for workloads that retrieve document context. Enter corpus size, chunk size, overlap, and Top K; the calculated retrieved tokens are added to each prompt.

  4. Confirm model pricing

    In Pricing, choose a currency and deployment scope, then add the models you plan to use. Enter PAYG and PTU values manually. Refresh official data can update supported PAYG prices, but currency changes never convert existing values and PTU fields may remain manual.

  5. Review forecast and routing

    In Forecast, choose a primary model for a single-model estimate or select tiered routing to use each scenario's model and allocation. Review monthly token demand, PAYG cost, and the 36-month compound-growth forecast.

  6. Assess capacity and share results

    Use Capacity to review peak TPM, required and provisioned PTUs, cost comparisons, and break-even timing. Export the model JSON for backup or reuse, CSV for analysis, and Print / PDF for a summary.

Before making a decision: validate model availability, prices, throughput assumptions, traffic shape, and commercial terms against current Azure sources.