You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取Word文档中加粗选项的完整列表项

问题描述

我有一份包含如下内容的MS Word(.docx)文档:

  1. best animal is :
    a. cat
    b. dog
    c. snake
  2. second best animal is:
    a. rhino
    b. tiger
    c. puma

其中dog和puma所在的整行是文档中的加粗文本(正确答案)。我尝试用Python遍历文档内容,识别加粗段落并将其作为正确答案输出为JSON,但目前仅能提取到加粗的文本内容(如dog、puma),无法提取对应的列表编号(如b.、c.),即无法输出“b. dog”“c. puma”这样的完整内容。

当前使用的代码:

import docx
import jsonpickle

doc_file = docx.Document('C:/Users/Admin/Downloads/test.docx')
json_output = []
for paragraph in doc_file.paragraphs:
    if paragraph.style.name.startswith('Heading 1') and paragraph.style.font.bold:
        json_output.append({'text': paragraph.text, 'bold': True})
   

json_string = jsonpickle.encode(json_output)

print(json_string)
解决方案

你的核心问题在于两个错误:

  1. 错误将判断条件限定为Heading 1样式,但正确选项的段落并非标题样式
  2. 仅检查段落样式的加粗属性,未考虑docx中按run(文本格式最小单元)存储的加粗格式

修改后的代码可提取完整的加粗段落文本(包含编号和内容):

import docx
import json

doc_file = docx.Document('C:/Users/Admin/Downloads/test.docx')
json_output = []

for paragraph in doc_file.paragraphs:
    # 遍历段落内所有run,判断是否存在加粗格式
    is_bold = False
    for run in paragraph.runs:
        if run.bold:
            is_bold = True
            break
    # 仅保留非空的加粗段落
    if is_bold and paragraph.text.strip():
        json_output.append({'text': paragraph.text.strip(), 'bold': True})

# 生成规范易读的JSON输出
json_string = json.dumps(json_output, ensure_ascii=False, indent=2)
print(json_string)
代码说明
  • 通过遍历段落内的run单元,准确识别加粗格式(docx中格式通常绑定到run而非整个段落)
  • 过滤空段落,避免无效内容混入结果
  • 使用标准json模块替代jsonpickle,输出更简洁规范的JSON结构

运行后输出的结果如下:

[
  {
    "text": "b. dog",
    "bold": true
  },
  {
    "text": "c. puma",
    "bold": true
  }
]

内容的提问来源于stack exchange,提问作者Alberto Valejo

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 11:05:14