bge-base-zh) is selected.readyParameter | Required | Description | Recommended Value |
Cache key policy | Yes | Cache matching scope | Select "Latest User Message" for single-turn Q&A, and select "Historical Dialogue Mode" for multi-turn conversations. |
TTL (seconds) | Yes | Cache validity period, ranging from 60-604800 seconds | 3600 (1 hour) |
Redis storage configuration | Yes | Cache storage address | Service address, port, username, and password |
Parameter | Required | Description | Recommended Value |
VDB service source | Yes | Select a pre-configured Tencent Cloud VDB service source | VDB instance in the same region |
Similarity Threshold | Yes | The value ranges from 0-1. A similarity score higher than this value is considered a hit. | 0.85 |
Maximum Token distance | Yes | Semantic caching is not used when the total number of request tokens exceeds this value. | 50 |
TTL (seconds) | Yes | Cache validity period, ranging from 60-604800 seconds | 3600 (1 hour) |
Database | Yes | The system automatically pulls the list of databases under the VDB instance. | - |
Vector storage configuration | Yes | Cache storage address | Service address, port, username, and password |
Parameter | Description |
Isolate Cache by User | Disabled by default. Independent caches are generated for identical requests from different users to enhance privacy. |
Cache Successful Responses Only | Enabled by default. Only successful responses with a 200 status code are cached. Error requests are not cached. |
Return cache identifier | Disabled by default. The X-Cache: HIT/MISS identifier is added to the response Header. |
Filtering Condition | Description | Option |
Consumer group | Filter by consumer group (all is optional) | All / Specified consumer group |
Model API | Filter by model API (all is optional) | All / Specified model API |
Time Range | Select the time window for data display. | Today/ This Week/ This Month/ Last 7 Days/ Last 30 Days/ Custom |
Metric Value | Description |
Total number of requests | Total number of LLM requests received by the gateway within the selected time range |
Exact cache hit count | L1 exact cache hit count, which displays the actual hit count and its percentage. |
Semantic cache hit count | L2 semantic cache hit count, which displays the actual hit count and its percentage. |
Gateway cache hit rate | (Exact hits + Semantic hits) / Total requests × 100% |
Gateway cache cost savings | Estimated cost savings from cache hits that avoid LLM calls (calculated based on model unit price) |
Similarity Range | Description |
0.9 - 1.0 | A high similarity hit indicates that the user request is very close to the cached content. |
0.8 - 0.9 | A medium similarity hit indicates a request with similar semantics but different wording. |
0.7 - 0.8 | A low similarity hit. Pay attention to whether a mismatch occurs. |
0.6 - 0.7 | A low similarity hit. It is recommended to check whether a misjudgment occurs. |
0 - 0.6 | An extremely low similarity hit. There may be a risk of mismatch. |
Panel Name | Description |
Cache hit rate trend | Shows the trend of cache hit rate changes within the selected time range (by hour/day granularity). |
Cache type distribution | Displays the hit ratio of exact cache and semantic cache in a pie chart. |
Response time comparison | Compares the average response time of cache-hit requests and cache-miss requests, and displays the improvement ratio. |
Column Name | Description |
Consumer group | Consumer group to which a request belongs |
Cache hit quantity | Number of requests with cache hits (exact + semantic) |
Number of cache misses | Number of requests that missed the cache and directly called the LLM |
Hit rate | Cache hit count / (hit count + miss count) × 100% |
Cost savings | Estimated cost savings from cache hits (¥) |
Tokens saved | Estimated number of tokens saved by cache hits |
Average response time for cache hits | Average response time of cache-hit requests (ms) |
Average response time for cache misses | Average response time of cache-miss requests (ms) |
Performance improvement ratio | Cache-miss average response time / cache-hit average response time |
Was this page helpful?
You can also Contact sales or Submit a Ticket for help.
Help us improve! Rate your documentation experience in 5 mins.
Feedback