You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python向AI API传输用户输入时如何防止敏感数据泄露?

Python 敏感数据脱敏/掩码处理方案(发送外部AI API前)

一、正则表达式掩码(基础可靠方案)

正则表达式是处理敏感数据掩码的基础方法,只要覆盖业务场景中常见的敏感数据模式,就能实现可靠的脱敏。以下是针对邮箱、手机号、身份证号的示例实现:

import re
import requests

def mask_sensitive_data(text):
    # 掩码邮箱地址(替换为***@***.***)
    text = re.sub(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}', '***@***.***', text)
    # 掩码国内手机号(保留区号和首尾部分,中间替换为*******)
    text = re.sub(r'(\+86)?1([3-9]\d{2})\d{4}(\d{4})', r'\1\2*******\3', text)
    # 掩码身份证号(全部替换为星号)
    text = re.sub(r'\d{17}[\dXx]', '*****************', text)
    return text

# 改造原API调用函数
def send_to_api(user_input):
    masked_input = mask_sensitive_data(user_input)
    response = requests.post(
        "https://api.example.com/analyze",
        json={"text": masked_input}
    )
    return response.json()

使用正则的注意事项:

  • 需根据业务场景调整正则规则(比如适配不同国家的手机号格式、带特殊符号的邮箱)
  • 用边缘测试用例验证(比如带空格的手机号、多级域名的邮箱)
  • 可逐步扩展规则覆盖更多敏感数据类型(如银行卡号、地址)

二、专业脱敏库(更优实践)

对于复杂场景,使用专业的脱敏库能提升识别准确率和开发效率,以下是两个常用库的实现:

1. Presidio(微软开源的敏感数据识别与脱敏工具)

Presidio支持多语言,内置数十种敏感实体识别规则(邮箱、手机号、身份证、信用卡号等),还可自定义实体类型。

首先安装依赖:

pip install presidio-analyzer presidio-anonymizer

示例代码:

import requests
from presidio_analyzer import AnalyzerEngine
from presidio_anonymizer import AnonymizerEngine
from presidio_anonymizer.entities import OperatorConfig

def anonymize_text(text):
    # 初始化分析器和脱敏引擎
    analyzer = AnalyzerEngine()
    anonymizer = AnonymizerEngine()
    
    # 识别文本中的敏感实体
    analysis_results = analyzer.analyze(text=text, language='zh')
    
    # 替换敏感数据为<REDACTED>,也可自定义替换规则
    anonymized_result = anonymizer.anonymize(
        text=text,
        analyzer_results=analysis_results,
        operators={
            "DEFAULT": OperatorConfig("replace", {"new_value": "<REDACTED>"})
        }
    )
    return anonymized_result.text

# 改造原API调用函数
def send_to_api(user_input):
    anonymized_input = anonymize_text(user_input)
    response = requests.post(
        "https://api.example.com/analyze",
        json={"text": anonymized_input}
    )
    return response.json()

2. Faker(生成假数据替换敏感信息)

如果需要保留数据格式但替换为虚假信息(比如外部API需要识别数据类型但不能传真实数据),可以用Faker生成符合格式的假数据。

首先安装依赖:

pip install faker

示例代码:

import re
import requests
from faker import Faker

fake = Faker('zh_CN')

def replace_sensitive_data(text):
    # 替换真实邮箱为假邮箱
    def replace_email(match):
        return fake.email()
    text = re.sub(r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}', replace_email, text)
    
    # 替换真实手机号为假手机号
    def replace_phone(match):
        prefix = match.group(1) if match.group(1) else ''
        return prefix + fake.phone_number()
    text = re.sub(r'(\+86)?1[3-9]\d{9}', replace_phone, text)
    
    return text

# 改造原API调用函数
def send_to_api(user_input):
    replaced_input = replace_sensitive_data(user_input)
    response = requests.post(
        "https://api.example.com/analyze",
        json={"text": replaced_input}
    )
    return response.json()

三、最佳实践补充

  • 明确敏感数据范围:先梳理业务中需要保护的敏感数据类型(如邮箱、手机号、身份证、银行卡号等),针对性制定脱敏规则
  • 测试覆盖:用包含各种边缘情况的测试用例验证脱敏效果,避免遗漏敏感数据
  • 避免过度脱敏:不要误处理非敏感数据(比如普通数字串不要当成手机号)
  • 合规适配:根据所在地区的隐私法规(如国内《个人信息保护法》、GDPR)调整脱敏策略,确保符合合规要求
  • 日志规范:脱敏后的数据可记录,但不要存储原始敏感数据(除非有合规要求的存储流程)

内容的提问来源于stack exchange,提问作者Rom Questa AI

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.01 17:52:27