如何用Python调用GCP DLP API对Word文档进行脱敏/去标识
Cloud DLP处理Word文档脱敏方案
原生支持情况
Cloud DLP原生支持.docx格式的Word文档脱敏,无需手动提取文本再处理。deidentify_content API不仅支持纯文本输入,还能直接解析处理包括Office文档在内的多种非结构化文件类型,内部会自动提取文档中的文本内容、应用脱敏规则后重新生成格式保持一致的文档。
实现步骤与代码示例
以下是针对.docx文档进行敏感数据脱敏的Python实现,以掩码处理邮箱和手机号为例:
def redact_word_document( project_id, input_filename, output_filename ): import google.cloud.dlp # 初始化DLP客户端 dlp_client = google.cloud.dlp_v2.DlpServiceClient() # 定义要脱敏的敏感信息类型 info_types = [{"name": "EMAIL_ADDRESS"}, {"name": "PHONE_NUMBER"}] # 构建去标识配置:用掩码替换敏感数据 deidentify_config = { "info_type_transformations": { "transformations": [ { "info_types": info_types, "primitive_transformation": { "character_mask_config": { "masking_character": "*", "number_to_mask": 100 # 掩码覆盖整个敏感内容 } } } ] } } # 读取Word文档字节数据 with open(input_filename, "rb") as f: file_data = f.read() # 构建待处理的内容项,指定文件类型为WORD item = { "byte_item": { "type_": google.cloud.dlp_v2.FileType.WORD, "data": file_data } } # 构建请求参数 parent = f"projects/{project_id}" request = { "parent": parent, "deidentify_config": deidentify_config, "item": item } # 调用DLP API response = dlp_client.deidentify_content(request=request) # 保存处理后的文档 with open(output_filename, "wb") as f: f.write(response.item.byte_item.data) print(f"已将脱敏后的文档保存至 {output_filename}")
代码说明
- 文件类型指定:通过
FileType.WORD告诉DLP解析.docx格式文档,API会保留原文档的格式(比如段落、字体、表格等),仅替换敏感内容。 - 脱敏规则配置:示例中使用
character_mask_config对邮箱和手机号进行掩码处理,你也可以根据需求替换为其他规则(比如替换为固定字符串、生成假名等)。 - 返回结果:
response.item.byte_item.data是处理后的文档字节数据,直接写入文件即可得到格式完整的脱敏Word文档。
针对旧版.doc格式的处理方案
如果需要处理老式的.doc格式文档,DLP不原生支持,可采用以下优雅方案:
- 先将.doc文件转换为.docx格式(可使用
libreoffice命令行工具或Python库如pywin32在Windows环境下转换)。 - 再使用上述
deidentify_contentAPI处理转换后的.docx文档。
内容的提问来源于stack exchange,提问作者Akhil Kv
相关产品推荐
相关产品推荐

