Scenarios
The AI Gateway supports data masking for sensitive information during both logging and request forwarding, helping you meet data security and compliance requirements.
Log Desensitization: Sensitive information is masked before the Prompt/Completion content is written to the LLM Log. This feature is applicable for preventing log leakage scenarios and does not affect the LLM's normal reception of raw data.
Forwarding Desensitization: Sensitive data is replaced with placeholders before the request is forwarded to the backend LLM vendor. This feature is applicable for scenarios with stringent data sovereignty and compliance requirements, preventing sensitive data from entering third-party model vendors.
|
Processing Timing | Before the LLM Log is written to | Before the LLM vendor is forwarded to |
Impact on LLM or Not | No impact. The LLM still receives the original data. | Yes, the LLM receives placeholders. |
Dependency on Packet Body Collection | Enable packet body collection. | Not required |
Processing Result | Masking (for example, 138****8888) | Placeholder (for example, [Phone Number]) |
Scenario | Most customers, ensuring compliance without affecting model performance. | Customers with stringent data sovereignty compliance requirements |
Note:
The DMask feature must be used in conjunction with the Request Body Collection switch. If body collection is not enabled, log desensitization does not take effect. Forwarding desensitization does not depend on the body collection switch.
Prerequisites
1. The model API has been created and is in a running state.
2. For log desensitization, you have enabled Request Body Collection (Prompt) or Response Body Collection (Completion) on the Model API Details page.
3. The sensitive data types that require desensitization (such as phone numbers, ID card numbers, and so on) have been identified.
Scenario 1: Configuring Log Masking Rules
Log desensitization masks sensitive data before it is written to the LLM Log. This feature is applicable for scenarios that require logging content while protecting sensitive information.
Step 1: Go to the Model API Details Page
2. On the instance list page, click the ID of the gateway instance you want to configure to go to its basic information page.
3. In the left sidebar, click Model Management, and then click the Model API tab.
4. Click the target model API name to go to the API details page, and then switch to the Basic Information Tab.
Step 2: Enable Package Collection
In the Log Configuration area, confirm that body collection is enabled.
Attention:
After body collection is enabled, the text content of requests and responses is written to log storage. Please confirm that the content does not contain user personal privacy data and complies with corporate compliance requirements.
Step 3: Configure Log Masking Rules
In the Log Desensitization Rules (Effective when body collection is enabled) area, click Configure Rules to go to the configuration dialog.
Configuration Parameter Description:
|
Enable or Not | Yes | The master switch for log desensitization, when the master switch is disabled, all desensitization rules become inactive. |
Built-in data masking types | No | Select the built-in data types that require masking. |
Data masking scope | No | Select the scope for data masking: Prompt only, Completion only, or both. |
Custom data masking rules | No | Add business-specific sensitive data identification rules. |
Built-in Desensitization Types:
|
Mobile Number | Mainland China 11-digit mobile phone number (1[3-9]\\d{9}) | 138****8888
|
ID Card Number | 18-digit or 15-digit ID card number | 110101****1234
|
Bank Card Number | 16- to 19-digit number (standard bank card format) | 6222****0001
|
Email | Standard email format | abc***@***.com
|
IP address | IPv4 Format | 192.168.*.*
|
Name | 2-4 Chinese characters | Zhang**
|
Step 4: Configure Custom Masking Rules (Optional)
If the built-in desensitization types cannot meet your business requirements, you can add custom rules:
1. Click Add Rule and enter the rule parameters:
|
Rule name | The name used to identify a rule. Up to 60 characters. | Employee ID
|
Matching regular expression | The regular expression used to match sensitive data. | E\\d{6}
|
Mask format | Mask format after desensitization | E***
|
Enabling log access | Whether to enable the rule | Enable by selecting |
Rule Name Restrictions:
Maximum 60 characters
Supports uppercase and lowercase letters in Chinese and English, numbers, and separators (-, _).
It cannot start with a number or a separator.
It cannot end with a separator.
2. Click OK to save the custom rule.
Step 5: Save the Configuration
Click OK in the configuration dialog, and the log desensitization rule takes effect immediately.
Scenario 2: Configure Forwarding Masking Rules
Forwarding Desensitization replaces sensitive data with placeholders before the request is forwarded to the LLM vendor. It is applicable for scenarios with stringent data sovereignty and compliance requirements.
Attention:
After forwarding desensitization is enabled, the LLM receives placeholders instead of real data, which may affect the model's understanding of the context. It is recommended to enable this feature only when there are explicit data sovereignty and compliance requirements.
Step 1: Go to the Model API Details Page
Same as Step 1 in Scenario 1.
Step 2: Configure Forwarding Masking Rules
In the Data Security Configuration (Forwarding Desensitization) area, click Configure Rules to go to the configuration dialog.
Configuration Parameter Description:
|
Enable or Not | Yes | The master switch for forwarding desensitization |
Built-in data masking types | No | Select the data types that require masking during forwarding. |
Placeholder format | Yes | The format of the placeholder after replacement. It is automatically set to [{type}] by default, supports customization, and the system automatically replaces {type} with the name of the corresponding sensitive data type. |
Handling data masking failures | Yes | Behavior when an exception occurs during data masking rule processing |
Custom data masking rules | No | Add business-specific sensitive data identification rules. |
Desensitization Failure Handling Policy:
|
Reject the request and return 500 (Recommended). | Reject the request directly when a data masking exception occurs to ensure no data leakage. | High-Compliance Scenarios |
Skip desensitization and forward directly | Forward the request as-is when data masking fails, without blocking it. | Not recommended, as it poses a data leakage risk. |
Step 3: Save the Configuration
Click OK in the configuration dialog, and the forwarding desensitization rule takes effect immediately.
Must-Knows
1. Body Collection Dependency: Log desensitization takes effect only when body collection is enabled. If body collection is not enabled, Prompt/Completion content is not recorded in the logs, and desensitization is not required.
2. Forwarding Desensitization Trade-offs: Forwarding desensitization alters the actual content received by the LLM, which may affect the quality of the model's output. It is recommended to enable this feature only in scenarios with stringent compliance requirements.
3. Performance Impact: Desensitization processing increases the CPU overhead of the gateway. Performance testing is recommended in high-concurrency scenarios.
4. Custom Regex Security: Exercise caution when writing regular expressions for custom desensitization rules to avoid overly broad rules causing legitimate content to be inadvertently desensitized.
5. Desensitization Scope Selection: It is recommended to select the desensitization scope based on your actual requirements:
Audit only the request input → Prompt only.
Audit only the model output → Completion only.
Require full audit → Prompt + Completion.
6. Caution for Name Recognition: Name desensitization has a relatively high false positive rate, as common Chinese words may be incorrectly identified as names. It is recommended to carefully evaluate before this feature is enabled.
7. Impact of Placeholders on the Model: Placeholders after forwarding desensitization (for example, [Phone Number]) are treated as plain text by the LLM, and the model cannot understand their actual meaning.