Feature Overview
Model pricing configuration is the foundational data module for AI Gateway cost management, used to configure the billing rates for model invocations. After configuration, AI Gateway calculates the cost in real time at the end of each large model invocation. This calculation is based on the three rate segments you configured—input, cache-hit input, and output—combined with the actual number of Tokens consumed by the request. This process then drives the cost data presentation in subsequent Token Consumption Statistics and Cost Analysis Reports.
It supports automatic synchronization of the latest official prices from 10 vendors. The official current prices can optionally be displayed. The synchronization cycle is 24 hours. It also supports batch updating of model unit prices using the official current prices.
Page Entrance
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, click Model Pricing Configuration in the Cost Management module.
Field Descriptions for the Configuration List
|
Model Vendor | The model provider, which supports selecting standard providers (such as Hunyuan, Google-Gemini, and so on), MaaS providers (such as TokenHub, AWS Bedrock, and GCP Vertex AI, and so on), and custom providers (requires selecting or manually entering the custom provider name). |
|
Model Name | Model name, such as hunyuan-t1 and gemini-2.5-flash |
|
Pricing Type | Fixed Price | Input Price: The unit price for input Tokens, for example, ¥1.0000/1M Tokens. Output Price: The unit price for output Tokens, for example, ¥4.0000/1M Tokens. Cache Input Price: The unit price for cache-hit input. - is displayed when it is not configured. |
| Tiered Pricing | Minimum Tokens (Inclusive): The minimum number of Tokens included in this tier. Maximum Tokens (Inclusive): The maximum number of Tokens included in this tier. (For the last tier, it is fixed as "No Limit".) Input Price: The unit price for input Tokens, for example, ¥1.0000/1M Tokens. Output Price: The unit price for output Tokens, for example, ¥4.0000/1M Tokens. Cache Input Price: The unit price for cache-hit input. - is displayed when it is not configured. Configure at least two tiers. You can configure up to five tiers of pricing by clicking Add Tier. The tier intervals must be continuous and non-overlapping. |
Measurement unit | Optional 1M Tokens/1K Tokens |
|
Currency | Currency unit, fixed as CNY |
|
Operation Steps
Adding a Model Unit Price Configuration
1. On the Model Pricing Configuration page, click Add Configuration in the upper-right corner. The Add Model Pricing Configuration dialog box will pop up.
2. Configure the following parameters:
|
Model Vendor | Yes | Select from a dropdown list. Options include mainstream providers preset by the system, such as OpenAI, DeepSeek, Qwen, Google-Gemini, and Hunyuan, as well as custom providers created by users in "Model Management". |
Model Name | Yes | Select from a dropdown list. Options are dynamically loaded based on the selected provider. |
Pricing Type | Yes | Select by clicking. Options: Fixed Price/Tiered Pricing. |
Measurement unit | Yes | Currently 1M Tokens (per million tokens) |
Currency | Yes | Fixed as CNY (Chinese Yuan) |
Input unit price. | Yes | Unit price for input tokens, measured in CNY/1M tokens. It is recommended to retain 4 decimal places. |
Output unit price. | Yes | Unit price for output tokens, measured in CNY/1M tokens. It is recommended to retain 4 decimal places. |
Cached input unit price | No | Unit price for cache-hit input tokens, measured in CNY/1M tokens. If left blank, the cost calculation automatically falls back to the input unit price. |
Minimum tokens (inclusive) | Yes (Tiered Pricing Only) | The minimum number of tokens included in this tier. |
Maximum tokens (inclusive) | Yes (Tiered Pricing Only) | The maximum number of tokens included in this tier (the last tier is fixed as "unlimited"). |
3. Click OK. The newly added configuration will appear in the list.
Batch Importing Model Unit Price Configurations
1. Click Batch Import. If this is your first time performing this operation, you need to first click Download Import Template in the pop-up window. After completing the template, upload it.
2. The template includes headers and sample data. The fields to be filled in the template must match those in the Add Single Model Pricing Configuration. You can view specific filling examples within the template.
3. Click Start Upload.
4. Click Parse File. The corresponding model pricing configurations will be parsed based on the filled template and displayed in a list for confirmation.
5. Click Confirm Import. Upon successful import, the above batch configurations will be automatically added to the configuration list. If any items fail, you can click View Details to view the details of the batch import task.
Editing a Model Unit Price Configuration
1. Locate the target configuration in the configuration list.
2. Click Edit in the Actions column. The Edit dialog box will pop up.
3. Click OK to save the modifications.
Exporting Model Unit Price Configurations
Click the Download button above the list to export all current model pricing configurations with one click.
Deleting a Model Unit Price Configuration
1. Locate the target configuration in the configuration list.
2. Click Delete in the Actions column.
3. In the pop-up confirmation dialog box, view the prompt: "Are you sure you want to delete the model pricing configuration? Once deleted, it cannot be recovered. Please confirm the operation.".
4. Click Confirm to complete the deletion.
Attention:
After deletion, the cost for this model cannot be calculated. In the related Token consumption statistics and cost analysis reports, the Cost column for this model will be displayed as - (which differs from ¥0.00, indicating "unit price not configured" rather than "zero cost").
Must-Knows
Impact of Unconfigured Unit Price: If the "model vendor + model name" corresponding to a request has no unit price configured, the cost field for that request is set to null and is displayed as - in the Token consumption statistics and cost analysis reports.
Custom Models: For private models or non-mainstream vendors, first create the corresponding model service in "Model Management". Ensure that the provider and model returned by the model service exactly match the unit price configuration.
Currency and Unit of Measurement: The currency is currently fixed to CNY, and the unit of measurement is fixed to 1M Tokens. USD or 1K Tokens are not currently supported.
Pricing Type: Only fixed pricing is currently supported. Other pricing types, such as tiered pricing, are planned for future versions.
FAQs
Q1: Why Are Some Line Items in My Cost Analysis Report Displayed as -?
A: Indicates that the unit price for this "model vendor + model name" is not configured. Go to the "Model Unit Price Configuration" page to add the corresponding unit price. New requests will then immediately start calculating costs. - indicates that the unit price is not configured, which has a different meaning from ¥0.00 (where the actual cost is zero).
Q2: What Happens If the Cache Input Unit Price Is Left Blank?
A: During cost calculation, the cache-hit portion is automatically calculated based on the Input Price. This does not affect the overall accuracy of billing, but you will not benefit from the cache discounts offered by model vendors (for example, vendors like DeepSeek and Anthropic offer significant price reductions for cache hits). It is recommended to consult the official vendor pricing and configure it accurately.
Q3: Can Multiple Unit Prices Be Configured for the Same Model?
A: No. Under the same gateway instance, the combination of "model vendor + model name" must be unique, and only one unit price configuration can exist. To adjust the price, use the Edit operation.
Q4: Will the Cost of Historical Calls Be Recalculated After the Unit Price Is Modified?
A: No. Price adjustments only apply to newly incoming requests. If you need to recalculate historical cost data, record the price adjustment timestamp in the remarks and perform the calculation offline after exporting the CSV.
Q5: Why Is the Option I Need Not Available in the Model Vendor and Model Name Drop-Down Lists?
A: The dropdown options are automatically populated based on the provider/model of the model services already created under Model Management > Model Service in the current gateway instance. First, create the corresponding model service in Model Management. The model option will then be visible on the unit price configuration page.