tencent cloud

Cloud Native Intelligent Gateway

Viewing Request Monitoring

다운로드
포커스 모드
폰트 크기
마지막 업데이트 시간: 2026-09-22 18:51:11
AI 번역

Scenarios

AI Gateway provides multi-dimensional monitoring metrics for running gateway instances to comprehensively monitor instance health and AI invocation quality. These metrics cover general gateway performance metrics (such as number of requests, latency, and error codes) and LLM-specific metrics for large model scenarios (such as Token consumption and model response time).
You can use these metrics to monitor the health of gateway instances and individual model APIs in real time, gain insights into AI invocation costs and performance, promptly address potential risks, and ensure the stability and cost controllability of your AI services. This document describes how to view gateway request monitoring through the TSF console.

Operation Steps

1. Log in to the Microservices Platform console. In the left sidebar, select AI Gateway to go to the instance list.
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, click Data Observation.
4. Click the Request Monitoring tag at the top of the page.

Supported Monitoring Metrics and Their Meanings

Instance/Node

These metrics apply to all traffic passing through the gateway and are used to evaluate the general performance and health of the gateway and backend services.
Metric Name
Metric Meaning
Total number of requests
Total number of requests, summed based on the selected time granularity
Average request latency
Average request latency, calculated as the average based on the selected time granularity.
Maximum request latency
Maximum request latency, calculated as the maximum value based on the selected time granularity.
Number of requests directly returned by the gateway
Number of requests that are not forwarded to the backend but directly responded to by the gateway (for example, when authentication fails or throttling is triggered), summed based on the selected time granularity.
Average gateway latency
Average time taken by the gateway itself to process requests.
Maximum gateway latency
Maximum time taken by the gateway itself to process requests.
Number of 2xx requests
Number of successful requests (for example, 200 OK) sent from the client to the AI gateway, summed based on the selected time granularity.
Number of 3xx requests
Number of requests redirected after they are sent from the client to the AI gateway, summed based on the selected time granularity.
Number of 4xx requests
Number of client errors directly returned by the gateway for illegal requests sent from the client to the AI gateway (for example, due to authentication failure or exceeding the throttling limit), such as 401 authentication failure, 403 insufficient permissions, and 429 throttling, summed based on the selected time granularity.
Number of 5xx requests
Number of server-side errors returned by the backend service after the AI gateway forwards messages to it (such as 500 backend exception, 502 backend invalid response, and 504 backend unreachable), summed based on the selected time granularity.
Number of 403 requests
Permission error. The backend is capable of processing the request but denies authorization access.
Number of 404 requests
Number of requests that failed to reach the backend service because the requested resource was not found on the backend server, summed based on the selected time granularity.
Number of 429 requests
Number of requests failed to be sent to the backend service because the requests are throttled, summed based on the selected time granularity
Number of 499 requests
Number of requests that failed to reach the backend service because the client actively disconnected before the backend responded, summed based on the selected time granularity.
Number of 502 requests
Number of errors where the gateway receives an invalid response from the backend server (usually due to a connection failure) while attempting to execute a backend request, summed based on the selected time granularity.
Number of 504 requests
Number of errors where the backend machine is unreachable when the gateway attempts to execute a backend request, summed based on the selected time granularity.
Number of requests forwarded to the backend by the gateway
Number of requests successfully forwarded by the gateway to the backend service, summed based on the selected time granularity.
Average backend latency
Average time taken by the backend service to process requests, calculated as the average based on the selected time granularity.
Maximum backend latency
Maximum time taken by the backend service to process requests, calculated as the maximum value based on the selected time granularity.
Number of backend 2xx requests
Number of successful requests (for example, 200 OK) processed by the backend service, summed based on the selected time granularity.
Number of backend 3xx requests
Number of redirection requests processed by the backend service, summed based on the selected time granularity.
Number of backend 4xx requests
Number of illegal requests sent to the backend service, summed based on the selected time granularity.
Number of backend 5xx requests
Number of server-side errors returned by the backend service (for example, 500 backend exception, 502 backend invalid response, 504 backend unreachable), summed based on the selected time granularity.
Number of backend 404 requests
Number of errors where the requested backend service resource is not found on the backend server, summed based on the selected time granularity.
Number of backend 429 requests
Number of errors where the backend service request fails because the request is throttled, summed based on the selected time granularity.
Number of backend 499 requests
Number of errors where the backend service request fails because the client actively disconnects before the backend responds, summed based on the selected time granularity.
Number of backend 502 requests
Number of errors where the backend service request fails because the backend service receives an invalid response, summed based on the selected time granularity.
Number of backend 504 requests
Number of errors where the backend service request fails because the backend machine is unreachable, summed based on the selected time granularity.

Model API

These metrics are specifically designed to monitor large model invocation scenarios, helping you analyze Token consumption costs and model provider performance.
Metric Name
Metric Meaning
Number of LLM HTTP requests
The number of HTTP calls initiated by the gateway to the large model provider. This metric directly reflects the invocation frequency of the model API.
Total tokens consumed by LLM
The total number of tokens consumed by the gateway from the large model provider, which is the sum of the actual tokens consumed for input (Prompt) and output (Completion). It is used to evaluate the total data throughput of Token consumption.
Tokens consumed by LLM prompt
The total number of tokens consumed by the model for the input (Prompt) part when the large model processes a request.
Tokens consumed by LLM completion
The total number of tokens consumed by the model for the output (Completion) part when the large model generates a response. This metric is one of the core bases for evaluating model invocation costs.
Average latency per LLM request (ms)
The average duration from when the gateway sends a request to the model provider to when it receives the complete response. This metric reflects the end-to-end response performance of the model provider.
Average latency per token by LLM provider (ms)
The average time spent by the model provider to consume each Token. This metric reflects the Token consumption speed of the model provider.

MCP

These metrics are used to monitor the health of MCP services. On this page, you can view the invocation details of each MCP service, including the number of requests, average latency, and request success rate, to quickly identify exceptions and perform capacity assessment.
Metric Name
Metric Meaning
MCP request quantity
Total number of requests received by the MCP service within the selected time range. This metric directly reflects the invocation frequency of the MCP Server.
Average MCP request latency (ms)
Average time taken by the MCP service to process requests (in milliseconds). This metric reflects the performance of the MCP Server.
MCP request success rate (%)
Percentage of successful requests among total requests invoked by the MCP service. This metric reflects the stability of the MCP Server.

도움말 및 지원

문제 해결에 도움이 되었나요?

피드백