You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python解析Word文档并提取含下划线单词至DataFrame的实现方案

Python解析Word文档并筛选含下划线的单词

依赖安装

先安装所需工具库:

pip install python-docx pandas

完整示例代码

import docx
import pandas as pd
import re

# 替换为你的目标Word文档路径
doc_path = "your_document.docx"

# 读取Word文档
doc = docx.Document(doc_path)

# 提取文档所有文本(含段落和表格)
all_text = []
# 提取段落文本
for para in doc.paragraphs:
    all_text.append(para.text)
# 提取表格文本(可选,按需启用)
for table in doc.tables:
    for row in table.rows:
        for cell in row.cells:
            all_text.append(cell.text)

# 合并文本并筛选含下划线的单词
combined_text = " ".join(all_text)
# 匹配包含至少一个下划线的完整单词,避免标点干扰
underscore_words = re.findall(r'\b\w*_\w*\b', combined_text)

# 转换为DataFrame并去重(去重步骤可选)
df = pd.DataFrame(underscore_words, columns=["含下划线的单词"])
df = df.drop_duplicates().reset_index(drop=True)

# 输出结果
print(df)

补充说明

  • 代码仅支持.docx格式文档,若需处理旧版.doc格式,Windows环境可使用pywin32库,跨平台可尝试textract库
  • 正则表达式\b\w*_\w*\b会匹配所有包含下划线的完整单词,包括以下划线开头/结尾的情况
  • 若不需要去除重复单词,可删除df = df.drop_duplicates().reset_index(drop=True)这一行

内容的提问来源于stack exchange,提问作者as_da_programming

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 21:10:28