You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Python提取PDF中非表格类结构化文本数据

PDF结构化文本提取故障排查请求

我需要从PDF中提取特定结构化文本,预期输出为包含word和type字段的结构化格式(对应预设示例样式)。先后尝试两种方案均未达成目标:

  • 使用Tabula库提取表格数据,输出结果不符合结构要求
  • 编写正则表达式匹配文本,无法正确拆分目标字段

以下是我编写的完整代码,恳请帮忙排查问题:

import tabula
import pandas as pd
import numpy as np


df = (pd.concat(
         tabula.read_pdf(
              "/content/drive/MyDrive/Stage/word.pdf", pages="all", pandas_options={"header": None}))
         .squeeze().str.extract(r"\)\s*([^\s]+)\s*([a-z\s,]+)?\s*([A-Z\s]+)?\s*(\w\d+)")
         .stack(dropna=False).strip().unstack()
         .set_axis(["word", "type", "comment", "suffix"], axis=1)
     [["word", "type"]] #uncomment this line to match your expected output
     )

df.to_excel("table.xlsx", index=False) #uncomment this line to make a spreadsheet
print(df)

内容的提问来源于stack exchange,提问作者tous

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 21:37:16