tencent cloud

Cloud Native Intelligent Gateway

Token Length Routing

Download
Focus Mode
Font Size
Last updated: 2026-09-22 18:43:15
AI-Translated

Scenarios

Token-length-based intelligent routing is a routing policy that automatically selects a more suitable model service based on the number of tokens in a request. After receiving a user request, the system automatically estimates the Token length of the request, matches the target model according to preset segmentation rules, and achieves an optimal balance between cost and performance.
Typical scenarios:
Multi-model cost optimization: The system integrates multiple models (GPT-4 with short context and high cost, Claude with medium context and medium cost, Gemini with long context and low cost) and automatically selects the most economical model based on the Token length of each request.
Long-short text routing: Short text requests (for example, 0-4K tokens) are routed to cost-effective models, while long text requests (for example, 32K+ tokens) are routed to long-context models, preventing truncation and cost waste.

Prerequisites

The AI gateway instance has been created and is in a running state.
Multiple model services have been configured to serve as a candidate service pool.
The Token length value for expected segmentation has been specified.

Operation Steps

Step 1: Go to the Model API Configuration Page

1. Log in to the Microservices Platform console. In the left sidebar, click AI Gateway > Instance List.
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. Choose Model Management > Model API in the left sidebar.
4. Click New or edit an existing API.
5. After completing the basic information configuration, go to Step 2: Select Model Service.

Step 2: Select the Service Type and Routing Policy

1. Select Service Type as Multi-Model Service.
2. If you need to filter the candidate service pool first, you can enable Tag Filtering (optional). After it is enabled, the system filters eligible model services based on the tags. If Tag Filtering is not enabled, the candidate service pool is the full list of model services.
Filtering Mode: Select AND (the service must meet all Tags) or OR (the service needs to meet any one Tag). AND is recommended.
Tag Configuration: Click Add Tag and enter the Tag's Key and Value.
3. In the Routing Policy section, select Token-Length Routing.

Step 3: Select the Tokenizer Encoding

Select the tokenization rule corresponding to your model. Different encodings directly affect the number of Tokens, API cost, and Chinese processing efficiency. Make your selection based on the model you actually use.
Code
Recommended Scenario
Feature
o200k_base (Recommended)
GPT-4o / o1 / o3 / GPT-5 series
The latest and most cost-effective. Offers the best support for Chinese, consuming fewer tokens for the same Chinese content.
cl100k_base​
GPT-4 / GPT-3.5 / Claude series
Universally compatible. Suitable for legacy OpenAI models or aligning Token counts with Claude.
p50k_base
Early GPT-3 / Codex
Biased towards code processing, with low efficiency for Chinese. Used only for legacy system maintenance.
r50k_base
Early GPT-2 / GPT-3
The oldest encoding, with extremely high token consumption for Chinese. Generally not recommended for selection.

Step 4: Configure the Default Target

This default target is used when Token counting fails, the rule is empty, or no segmentation rule is matched.
1. Select Secondary Routing Policy: Single-Model Service, Weighted Routing, or Latency-First Routing.
2. Complete the subsequent configuration based on the selected secondary routing policy.

Step 5: Configure the Segmentation Rules

1. In the Segmentation Rules section, configure multiple segments based on Token length intervals and set the following parameters for each segment:
Token Length Interval: [Minimum Tokens included in the segment, Maximum Tokens included in the segment]. The upper bound of the previous segment serves as the lower bound of the next segment. Segments must be contiguous, and intervals must not overlap.
Typical Configuration Example:
Short Text: [1, 100] / [101, 500] / [501, 2000]
Medium Text: [1, 1000] / [1001, 8000] / [8001, 32000]
Long Text: [1, 4000] / [4001, 32000] / [32001, 200000]
Secondary Routing Policy: Single-Model Service, Weighted Routing, or Latency-First Routing.
Complete the subsequent configuration based on the selected secondary routing policy. When multiple models are configured within a segment, they are automatically distributed among the models according to the selected policy.
2. You can add a new segment by clicking Add Segmentation Rule.

Step 6: Confirm and Save

Click OK to save the configuration. The intelligent routing policy takes effect after the configuration validity check is passed.



Help and Support

Was this page helpful?

Help us improve! Rate your documentation experience in 5 mins.

Feedback