You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python获取Word中指定关键词后的表格内容

通过关键词定位Word文档中目标表格的Python实现

要实现通过关键词定位其后续表格的需求,我们可以直接遍历Word文档的正文元素,跟踪关键词段落的位置,然后捕获紧随其后的第一个表格。以下是具体代码实现:

完整代码示例

from docx import Document
from docx.text.paragraph import Paragraph
from docx.table import Table

def get_table_after_keyword(doc_path, target_keyword):
    doc = Document(doc_path)
    found_keyword = False
    target_table = None

    # 遍历文档正文的所有元素(段落和表格)
    for element in doc.body._element:
        # 判断当前元素是否为段落
        if element.tag.endswith('p'):
            para = Paragraph(element, doc)
            if target_keyword in para.text:
                found_keyword = True
        # 若已找到关键词,且当前元素是表格,则记录并终止遍历
        elif element.tag.endswith('tbl') and found_keyword:
            target_table = Table(element, doc)
            break

    if not target_table:
        return None  # 未找到匹配的表格或关键词

    # 提取表格内容为二维列表
    table_content = []
    for row in target_table.rows:
        row_data = [cell.text.strip() for cell in row.cells]
        table_content.append(row_data)
    
    return table_content

# 使用示例
if __name__ == "__main__":
    result = get_table_after_keyword("your_document.docx", "目标关键词")
    if result:
        for row in result:
            print(row)
    else:
        print("未找到匹配的表格或关键词")

代码说明

  • 遍历混合元素:通过doc.body._element访问文档的底层XML元素,这样可以按顺序遍历段落和表格,解决了doc.paragraphs和doc.tables无法体现元素顺序的问题。
  • 关键词检测:遇到段落时检查是否包含目标关键词,标记found_keyword为True。
  • 捕获目标表格:标记后遇到的第一个表格即为目标表格,直接终止遍历提升效率。
  • 内容提取:将表格内容转换为二维列表,方便后续处理或输出。

注意事项

  • 若关键词出现多次,代码会捕获第一次出现后紧随的表格;如需处理所有匹配情况,可移除break并收集所有符合条件的表格。
  • 确保已安装python-docx库:pip install python-docx

内容的提问来源于stack exchange,提问作者TJJJ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 14:20:27