如何用Python提取Word文档中加粗选项的完整列表项
问题描述
我有一份包含如下内容的MS Word(.docx)文档:
- best animal is :
a. cat
b. dog
c. snake- second best animal is:
a. rhino
b. tiger
c. puma
其中dog和puma所在的整行是文档中的加粗文本(正确答案)。我尝试用Python遍历文档内容,识别加粗段落并将其作为正确答案输出为JSON,但目前仅能提取到加粗的文本内容(如dog、puma),无法提取对应的列表编号(如b.、c.),即无法输出“b. dog”“c. puma”这样的完整内容。
当前使用的代码:
import docx import jsonpickle doc_file = docx.Document('C:/Users/Admin/Downloads/test.docx') json_output = [] for paragraph in doc_file.paragraphs: if paragraph.style.name.startswith('Heading 1') and paragraph.style.font.bold: json_output.append({'text': paragraph.text, 'bold': True}) json_string = jsonpickle.encode(json_output) print(json_string)
解决方案
你的核心问题在于两个错误:
- 错误将判断条件限定为
Heading 1样式,但正确选项的段落并非标题样式 - 仅检查段落样式的加粗属性,未考虑docx中按
run(文本格式最小单元)存储的加粗格式
修改后的代码可提取完整的加粗段落文本(包含编号和内容):
import docx import json doc_file = docx.Document('C:/Users/Admin/Downloads/test.docx') json_output = [] for paragraph in doc_file.paragraphs: # 遍历段落内所有run,判断是否存在加粗格式 is_bold = False for run in paragraph.runs: if run.bold: is_bold = True break # 仅保留非空的加粗段落 if is_bold and paragraph.text.strip(): json_output.append({'text': paragraph.text.strip(), 'bold': True}) # 生成规范易读的JSON输出 json_string = json.dumps(json_output, ensure_ascii=False, indent=2) print(json_string)
代码说明
- 通过遍历段落内的
run单元,准确识别加粗格式(docx中格式通常绑定到run而非整个段落) - 过滤空段落,避免无效内容混入结果
- 使用标准
json模块替代jsonpickle,输出更简洁规范的JSON结构
运行后输出的结果如下:
[ { "text": "b. dog", "bold": true }, { "text": "c. puma", "bold": true } ]
内容的提问来源于stack exchange,提问作者Alberto Valejo
相关产品推荐
相关产品推荐

