调用pd.read_pdf报错module 'pandas' has no attribute 'read_pdf'如何解决
问题解决方案
错误根因
pandas官方原生库不提供read_pdf方法,该报错是调用了不存在的类属性导致的。你看到的包含pd.read_pdf的示例大概率是误用了其他第三方库的方法,或依赖了未安装的pandas扩展包。
不同格式文件的读取方案
你需要分别使用对应工具库读取PDF、DOCX、TXT格式文件,以下是可直接复用的方案:
读取PDF文件
优先选用专门处理PDF表格的第三方库,读取结果可直接转换为pandas DataFrame格式:
- 可选库1:camelot-py
安装命令:pip install camelot-py[cv]
调用示例:import camelot # pages参数指定读取的页码,all为读取全部页 tables = camelot.read_pdf('目标文件.pdf', pages='1') # 取第一个表格转为DataFrame child1Pdf = tables[0].df - 可选库2:tabula-py
安装命令:pip install tabula-py
调用示例:import tabula # 读取PDF中表格,直接返回DataFrame列表 dfs = tabula.read_pdf('目标文件.pdf', pages=1) child1Pdf = dfs[0]
读取DOCX文件
使用python-docx库解析DOCX内容:
安装命令:pip install python-docx
调用示例:
from docx import Document import pandas as pd doc = Document('目标文件.docx') # 提取全文本 full_text = '\n'.join([para.text for para in doc.paragraphs]) # 提取文档中第一个表格转为DataFrame table = doc.tables[0] table_data = [[cell.text for cell in row.cells] for row in table.rows] docx_df = pd.DataFrame(table_data[1:], columns=table_data[0])
读取TXT文件
pandas原生支持读取TXT文件,无需额外安装依赖:
import pandas as pd # 可按需指定encoding、sep分隔符等参数 txt_df = pd.read_table('目标文件.txt', encoding='utf-8')
原代码额外问题修正
你提供的示例代码中os.chdir(directory+f"\\{fileDir}")一行的directory变量未提前定义,会触发NameError,建议提前把根路径赋值给该变量,示例:
# 提前定义根目录变量 directory = "C:\\Users\\adity\\Documents\\Parent" os.chdir(directory)
内容的提问来源于stack exchange,提问作者Aditya Mishra
相关产品推荐
相关产品推荐

