如何用R从Word文档标题下提取高亮文本到DataFrame?
提取Word文档高亮文本并生成Excel表格的解决方案
核心思路
用python-docx库解析Word文档的内容与格式信息,精准提取PID、X、Y区块下的高亮文本,再通过openpyxl将结果写入Excel表格。
步骤1:安装依赖库
执行以下命令安装所需工具:
pip install python-docx openpyxl
步骤2:编写Python脚本
以下是可直接运行的批量处理脚本:
import os from docx import Document from openpyxl import Workbook def extract_highlighted_text(paragraph): """提取单个段落中的高亮文本""" highlighted_text = [] for run in paragraph.runs: # 识别所有带高亮的文本,若需指定颜色可替换为具体色值判断 if run.font.highlight_color is not None: highlighted_text.append(run.text.strip()) return ' '.join(highlighted_text) def process_docx(file_path): """处理单个docx文件,返回PID、X、Y对应的高亮内容""" doc = Document(file_path) current_section = None pid = "" x_content = [] y_content = [] for para in doc.paragraphs: para_text = para.text.strip() if not para_text: continue # 识别标题区块(可根据实际文档标题格式调整判断逻辑) if para_text == "PID": current_section = "PID" continue elif para_text == "X": current_section = "X" continue elif para_text == "Y": current_section = "Y" continue # 收集对应区块的内容 if current_section == "PID": pid = para_text # PID为6位标识符,直接提取整段文本 elif current_section == "X": highlight = extract_highlighted_text(para) if highlight: x_content.append(highlight) elif current_section == "Y": highlight = extract_highlighted_text(para) if highlight: y_content.append(highlight) return { "PID": pid, "X": '; '.join(x_content), "Y": '; '.join(y_content) } def batch_process_and_export(folder_path, output_excel): """批量处理文件夹下的docx文件,结果写入Excel""" wb = Workbook() ws = wb.active ws.append(["PID", "X", "Y"]) # 写入表头 # 遍历目标文件夹下所有docx文件 for filename in os.listdir(folder_path): if filename.endswith(".docx"): file_path = os.path.join(folder_path, filename) try: result = process_docx(file_path) ws.append([result["PID"], result["X"], result["Y"]]) print(f"处理完成:{filename}") except Exception as e: print(f"处理失败 {filename}:{str(e)}") wb.save(output_excel) print(f"结果已导出到:{output_excel}") # 示例调用(替换为你的文件夹路径和输出Excel路径) if __name__ == "__main__": target_folder = "./docx_files" # 存放docx文件的文件夹路径 output_file = "./highlight_result.xlsx" # 输出Excel的路径 batch_process_and_export(target_folder, output_file)
关键调整点
- 若文档中PID、X、Y标题有特殊格式(比如加粗、特定字号),可修改标题识别逻辑,例如通过
para.runs[0].font.bold判断是否为标题段落。 - 如需提取特定颜色的高亮(比如黄色),可将
run.font.highlight_color is not None改为run.font.highlight_color == 4(4对应黄色,其他颜色值可查阅python-docx的颜色枚举)。 - X/Y下的高亮文本默认用分号拼接,可根据需求修改分隔符。
内容的提问来源于stack exchange,提问作者user23518192
相关产品推荐
相关产品推荐

