使用gspread-dataframe将PDF数据导入Google Sheets时遇属性错误
解决AttributeError: 'list' object has no attribute 'shape'问题
错误根源:
你用tabula提取PDF数据时,multiple_tables='FALSE'传的是字符串类型,但tabula的read_pdf方法要求该参数为布尔值(True/False)。Python会将非空字符串视为真值,因此read_pdf会返回一个包含DataFrame的列表,而非单个DataFrame。后续调用set_with_dataframe时,传入的df是列表,而该方法需要的是pandas DataFrame,因此触发shape属性不存在的错误。修复方案:
修改PDF提取代码中的参数,将字符串形式的布尔值改为原生布尔值:import tabula import pandas as pd file_path = 'TnPresidentbyCountyNov2016.pdf' # 修正multiple_tables和stream参数为布尔值 df = tabula.read_pdf(file_path, pages='all', multiple_tables=False, stream=True)额外验证建议:
提取数据后可以添加一行代码确认df的类型,确保是pandas DataFrame:print(type(df)) # 正确输出应为 <class 'pandas.core.frame.DataFrame'>如果PDF中包含多个表格,你需要从返回的列表中取出对应表格的DataFrame(例如
df = df[0]获取第一个表格),再传入set_with_dataframe。
内容的提问来源于stack exchange,提问作者Wayne Shaw
相关产品推荐
相关产品推荐

