tencent cloud

Cloud Native Intelligent Gateway

Overview

Download
Focus Mode
Font Size
Last updated: 2026-09-22 18:43:14
AI-Translated & Reviewed

Scenarios

An intelligent routing policy is the traffic distribution capability of the AI Gateway in multi-model service scenarios. By configuring an intelligent routing policy, you can:
It enables load balancing and traffic distribution for multi-model services.
It automatically matches the corresponding backend service based on the model name.
It flexibly controls traffic ratios based on weight configurations.
It automatically routes requests to the optimal model based on the semantic content of user requests.
It isolates and routes traffic from different tenants/environments based on request parameters, such as headers.
It filters available services by service Tag to implement model filtering.
It automatically selects the service with the fastest response based on real-time latency.
It supports scenarios such as A/B testing and canary releases.
It achieves high availability by working with the Fallback mechanism.
This document guides you on how to configure and use intelligent routing policies in the AI Gateway console.

Prerequisites

The AI Gateway instance has been created.
At least one model service (either a built-in vendor service or a custom model service) has been created.
The model API has been created.

Routing Policy Overview

AI Gateway supports multiple intelligent routing policies:
Routing Policies
Scenario
Decision Basis
Applicable Scope
Route by model name
Precise matching for multi-model services
Match the model field in the request with the model name configured in the model service.
A single-model API needs to access the same model from different suppliers.
Weight-based routing
Load balancing, A/B testing, and canary release
Randomly allocate traffic according to the configured weight ratio.
Multi-instance deployment of the same model or hybrid deployment across multiple suppliers
Intent-based routing
Classify and distribute based on semantic content
Call an intent recognition model to analyze the semantics of the request, and then route based on the identified intent.
Multi-task model, service tiering by problem complexity
Parameter-based routing
Multi-tenant isolation, environment isolation, and user traffic distribution
Match routing rules based on the values of the request Header parameters.
Route traffic to different services based on tenant/environment/user type.
Tag-based routing
Filtering by model tags
Automatically filter available services through model service tags (Key-Value).
A large number of models require filtering by tags.
Latency-first routing
Cross-region multi-path routing and response experience optimization
Select the optimal path based on real-time latency monitoring data of the model service.
Multi-region deployment, multi-IDC access scenarios


Help and Support

Was this page helpful?

Help us improve! Rate your documentation experience in 5 mins.

Feedback