What is the pricing for the FW-Kimi-K3 (by Fireworks) model in Azure AI Foundry?

The Max Coder 25 Reputation points
2026-08-06T16:52:16.1666667+00:00

Here's a revised description that conveys frustration while staying professional enough for a public Microsoft Q&A:

I'm using the FW-Kimi-K3 model offered through Azure's partnership with Fireworks AI and it is deployed in Azure AI Foundry (East US) under a Microsoft for Startups credit subscription. I cannot find published pricing for it anywhere: not on the Azure pricing page, the Foundry model catalog, or the Retail Prices API. This is frankly unacceptable for a service that is actively billing customers.

From my Cost Management records, usage appears to be billed at approximately $10 per million tokens, and this same flat rate is applied to every token type which is input, cached input, and output with no cache-read discount whatsoever. On July 28 alone, my deployment logged 18.55M input tokens and 272.1K output tokens, and approximately $188 in startup credits were debited. Charging the same rate for cached input tokens as for output tokens is unreasonable and inconsistent with how every other provider prices cache reads.

What makes this worse is that there was no way for me to know this rate in advance. The pricing is not documented, the meter names in Cost Management are generic (Model 12, Model 14, Model 15), and there is no indication that a flat $10/M rate would apply across all token types. As a startup relying on credits to evaluate models, I had no opportunity to estimate or control costs before they were incurred.

Could someone from Microsoft please explain:

  1. The official per-million-token price for FW-Kimi-K3 (input, output, and cached input separately, if applicable), and whether it is set by Fireworks or Azure.
  2. Whether a cached-input discount exists for this model, and if not, why not.
  3. Where this pricing is published — or confirmation that it is currently undocumented and being billed without public disclosure.
  4. Why a flat rate is applied to all token types rather than differentiated input/output/cache pricing.

Startups on credit subscriptions deserve transparent, documented pricing before being billed. Without it, we cannot responsibly use or evaluate these models. Thank you.


Want me to tone the frustration up or down, or is this the right register?

Foundry Models
Foundry Models

A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference


Answer accepted by question author
Karnam Venkata Rajeswari 5,170 Reputation points Microsoft External Staff Moderator
2026-08-06T17:44:43.1533333+00:00

Hello @The Max Coder ,

Welcome to Microsoft Q&A .Thank you for reaching out to us.

The behavior being observed appears to stem from a difference between the published FW-Kimi-K3 pricing and the charges visible in Cost Management. While the model pricing is now publicly documented, the billing data currently available does not provide sufficient detail to determine how each token category was rated, which is why a meter-level reconciliation is required before drawing conclusions about the applied charges.

Based on the published documentation, FW-Kimi-K3 has separate pricing for each token type:

  • Input tokens: USD $3.30 per 1 million tokens
  • Cached input tokens: USD $0.33 per 1 million tokens
  • Output tokens: USD $16.50 per 1 million tokens

This confirms that FW-Kimi-K3 uses differentiated pricing and includes a documented cached-input discount. The published pricing therefore does not indicate a flat rate across all token types

Regarding pricing ownership, the Azure AI Foundry documentation states that model providers define licensing terms and pricing for partner-hosted serverless models, while Azure provides the hosting platform, deployment experience, governance, and billing infrastructure.

The difficulty locating pricing information is understandable. Although pricing is published, partner-hosted offerings are not always surfaced consistently across pricing tools such as the Azure Pricing Calculator, Retail Prices API, or other catalog experiences. Based on the currently available documentation, this appears to be a pricing discoverability issue rather than a situation where pricing is entirely undocumented

An additional consideration is the deployment billing route. The published Microsoft for Startups policy states that sponsorship credits apply to models sold and billed directly by Azure, while models billed through third-party providers, partner services, or Azure Marketplace are not eligible. Confirming whether the deployment is Azure-direct billed or Marketplace-billed may therefore help explain how sponsorship credits are being applied.

Please check if the following steps help-

  1. Confirm the exact model name, deployment type, region, and offer type from the deployment Properties page.
  2. Verify whether the deployment is Azure-direct billed or Marketplace-billed.
  3. Download the detailed usage and charges report from Cost Management.
  4. Review the records associated with Model 12, Model 14, and Model 15, including:
    • Meter ID
    • Meter Name
    • Quantity
    • Effective Price
    • Cost
    • Product Name
    • Resource ID
  5. Compare the Meter IDs against the applicable billing price sheet.
  6. Validate that Quantity × EffectivePrice reconciles with the meter-level charge, as documented in Cost Management guidance.
  7. Compare the billed quantities with the actual input, cached-input, and output token usage.

The following references might be helpful , please check them out

Please let us know if the response was helpful

Thank you

Was this answer helpful?

1 person found this answer helpful.
0 comments No comments

0 additional answers

Sort by: Most helpful

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.