You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python调用GCP DLP API对Word文档进行脱敏/去标识

Cloud DLP处理Word文档脱敏方案

原生支持情况

Cloud DLP原生支持.docx格式的Word文档脱敏,无需手动提取文本再处理。deidentify_content API不仅支持纯文本输入,还能直接解析处理包括Office文档在内的多种非结构化文件类型,内部会自动提取文档中的文本内容、应用脱敏规则后重新生成格式保持一致的文档。

实现步骤与代码示例

以下是针对.docx文档进行敏感数据脱敏的Python实现,以掩码处理邮箱和手机号为例:

def redact_word_document(
    project_id,
    input_filename,
    output_filename
):
    import google.cloud.dlp

    # 初始化DLP客户端
    dlp_client = google.cloud.dlp_v2.DlpServiceClient()

    # 定义要脱敏的敏感信息类型
    info_types = [{"name": "EMAIL_ADDRESS"}, {"name": "PHONE_NUMBER"}]

    # 构建去标识配置:用掩码替换敏感数据
    deidentify_config = {
        "info_type_transformations": {
            "transformations": [
                {
                    "info_types": info_types,
                    "primitive_transformation": {
                        "character_mask_config": {
                            "masking_character": "*",
                            "number_to_mask": 100  # 掩码覆盖整个敏感内容
                        }
                    }
                }
            ]
        }
    }

    # 读取Word文档字节数据
    with open(input_filename, "rb") as f:
        file_data = f.read()

    # 构建待处理的内容项,指定文件类型为WORD
    item = {
        "byte_item": {
            "type_": google.cloud.dlp_v2.FileType.WORD,
            "data": file_data
        }
    }

    # 构建请求参数
    parent = f"projects/{project_id}"
    request = {
        "parent": parent,
        "deidentify_config": deidentify_config,
        "item": item
    }

    # 调用DLP API
    response = dlp_client.deidentify_content(request=request)

    # 保存处理后的文档
    with open(output_filename, "wb") as f:
        f.write(response.item.byte_item.data)

    print(f"已将脱敏后的文档保存至 {output_filename}")

代码说明

  • 文件类型指定:通过FileType.WORD告诉DLP解析.docx格式文档,API会保留原文档的格式(比如段落、字体、表格等),仅替换敏感内容。
  • 脱敏规则配置:示例中使用character_mask_config对邮箱和手机号进行掩码处理,你也可以根据需求替换为其他规则(比如替换为固定字符串、生成假名等)。
  • 返回结果:response.item.byte_item.data是处理后的文档字节数据,直接写入文件即可得到格式完整的脱敏Word文档。

针对旧版.doc格式的处理方案

如果需要处理老式的.doc格式文档,DLP不原生支持,可采用以下优雅方案:

  • 先将.doc文件转换为.docx格式(可使用libreoffice命令行工具或Python库如pywin32在Windows环境下转换)。
  • 再使用上述deidentify_content API处理转换后的.docx文档。

内容的提问来源于stack exchange,提问作者Akhil Kv

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.22 04:45:30