You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python-docx提取Word文档红色文本及红文本间内容?

提取Word文档中的红色文本及间隔普通文本

需求:读取包含红色文字的Word文档,提取所有红色文本片段、红色文本之间的普通文本,分别存入对应列表(按你的示例,最终可拆分为「第一个红色文本」「红色间普通文本」「第二个红色文本」三个独立列表)

示例文档内容:

Here is an example of a paragraph with no red colored words. This is another example of a paragraph with no red colored text.
This is another example of a paragraph with no red colored text.
red colored words here:HRD:100 Some text between red colored text. [RED:TEXT:1111]
This is another example of a paragraph with no red colored text.
This is another example of a paragraph with no red colored text.

你现有代码仅能提取红色文本片段:

import docx
from docx.shared import RGBColor


def readtxt(filename):
    doc = docx.Document(filename)
    fullText = []
    for para in doc.paragraphs:
        for run in para.runs:
            if run.font.color.rgb == RGBColor(255, 000, 000):
                fullText.append(run.text)
    return fullText

fullText = readtxt('filepath.docx')

修改后的完整代码

import docx
from docx.shared import RGBColor


def extract_target_texts(filename):
    doc = docx.Document(filename)
    red_texts = []          # 存储所有红色文本片段
    interval_texts = []     # 存储红色文本之间的普通文本
    temp_normal = []        # 临时收集当前普通文本片段

    for para in doc.paragraphs:
        in_red_segment = False
        for run in para.runs:
            # 判断当前run是否为红色文本(兼容无颜色设置的情况)
            is_red = run.font.color.rgb == RGBColor(255, 0, 0) if run.font.color else False
            
            if is_red:
                # 如果之前在收集普通文本,先存入间隔列表
                if temp_normal:
                    interval_texts.append(''.join(temp_normal))
                    temp_normal = []
                red_texts.append(run.text)
                in_red_segment = True
            else:
                # 非红色文本,加入临时收集列表
                temp_normal.append(run.text)
                in_red_segment = False

    return red_texts, interval_texts

# 调用并拆分出你需要的三个列表
red_list, interval_list = extract_target_texts('filepath.docx')
# 按示例内容拆分
list_red1 = [red_list[0]] if len(red_list) >=1 else []
list_interval = [interval_list[0]] if len(interval_list)>=1 else []
list_red2 = [red_list[1]] if len(red_list)>=2 else []

代码说明

  1. 通过temp_normal临时缓存普通文本,遇到红色文本时自动将缓存的普通文本存入interval_texts
  2. 兼容了段落中无颜色设置的run,避免出现属性不存在的报错
  3. 最终可根据red_texts和interval_texts的内容,直接拆分出你需要的三个独立列表;如果文档中有多组红色文本,也能批量提取所有红色片段和它们之间的普通文本

内容的提问来源于stack exchange,提问作者user20395731

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 05:20:28