如何用Python-docx提取Word文档红色文本及红文本间内容?
提取Word文档中的红色文本及间隔普通文本
需求:读取包含红色文字的Word文档,提取所有红色文本片段、红色文本之间的普通文本,分别存入对应列表(按你的示例,最终可拆分为「第一个红色文本」「红色间普通文本」「第二个红色文本」三个独立列表)
示例文档内容:
Here is an example of a paragraph with no red colored words. This is another example of a paragraph with no red colored text.
This is another example of a paragraph with no red colored text.
red colored words here:HRD:100 Some text between red colored text. [RED:TEXT:1111]
This is another example of a paragraph with no red colored text.
This is another example of a paragraph with no red colored text.
你现有代码仅能提取红色文本片段:
import docx from docx.shared import RGBColor def readtxt(filename): doc = docx.Document(filename) fullText = [] for para in doc.paragraphs: for run in para.runs: if run.font.color.rgb == RGBColor(255, 000, 000): fullText.append(run.text) return fullText fullText = readtxt('filepath.docx')
修改后的完整代码
import docx from docx.shared import RGBColor def extract_target_texts(filename): doc = docx.Document(filename) red_texts = [] # 存储所有红色文本片段 interval_texts = [] # 存储红色文本之间的普通文本 temp_normal = [] # 临时收集当前普通文本片段 for para in doc.paragraphs: in_red_segment = False for run in para.runs: # 判断当前run是否为红色文本(兼容无颜色设置的情况) is_red = run.font.color.rgb == RGBColor(255, 0, 0) if run.font.color else False if is_red: # 如果之前在收集普通文本,先存入间隔列表 if temp_normal: interval_texts.append(''.join(temp_normal)) temp_normal = [] red_texts.append(run.text) in_red_segment = True else: # 非红色文本,加入临时收集列表 temp_normal.append(run.text) in_red_segment = False return red_texts, interval_texts # 调用并拆分出你需要的三个列表 red_list, interval_list = extract_target_texts('filepath.docx') # 按示例内容拆分 list_red1 = [red_list[0]] if len(red_list) >=1 else [] list_interval = [interval_list[0]] if len(interval_list)>=1 else [] list_red2 = [red_list[1]] if len(red_list)>=2 else []
代码说明
- 通过
temp_normal临时缓存普通文本,遇到红色文本时自动将缓存的普通文本存入interval_texts - 兼容了段落中无颜色设置的run,避免出现属性不存在的报错
- 最终可根据
red_texts和interval_texts的内容,直接拆分出你需要的三个独立列表;如果文档中有多组红色文本,也能批量提取所有红色片段和它们之间的普通文本
内容的提问来源于stack exchange,提问作者user20395731
相关产品推荐
相关产品推荐

