求助:使用Python提取PDF中非表格类结构化文本数据
PDF结构化文本提取故障排查请求
我需要从PDF中提取特定结构化文本,预期输出为包含word和type字段的结构化格式(对应预设示例样式)。先后尝试两种方案均未达成目标:
- 使用Tabula库提取表格数据,输出结果不符合结构要求
- 编写正则表达式匹配文本,无法正确拆分目标字段
以下是我编写的完整代码,恳请帮忙排查问题:
import tabula import pandas as pd import numpy as np df = (pd.concat( tabula.read_pdf( "/content/drive/MyDrive/Stage/word.pdf", pages="all", pandas_options={"header": None})) .squeeze().str.extract(r"\)\s*([^\s]+)\s*([a-z\s,]+)?\s*([A-Z\s]+)?\s*(\w\d+)") .stack(dropna=False).strip().unstack() .set_axis(["word", "type", "comment", "suffix"], axis=1) [["word", "type"]] #uncomment this line to match your expected output ) df.to_excel("table.xlsx", index=False) #uncomment this line to make a spreadsheet print(df)
内容的提问来源于stack exchange,提问作者tous
相关产品推荐
相关产品推荐

