Pandas报错:'Series' object has no attribute 'Columns' 数据聚合求助
问题解决:PDF表格空行聚合时的"Series object has no attribute 'columns'"错误
问题背景
需求:当表格A、B列为空时,将对应行的C列内容聚合到上方非空行的C列中。现有代码在部分PDF表格可正常运行,但处理某特定PDF表格时抛出错误:'Series' object has no attribute 'columns'。
原始代码
from tabula import read_pdf from tabulate import tabulate import pandas as pd import numpy as np Page_No = 45 tables = read_pdf('/content/210812154731_DECK CRANE - MACGREGOR HAGGLUND - TG 5795-36.5 130-2 - N00024-93-C-2220 - PARTS MANUAL.pdf', pages=Page_No) data_df = pd.DataFrame(tables[0]) out= data_df.ffill().groupby(['Item', 'SMR','CAGE','Part Number','Qty'], as_index=False)['Description'].agg(' '.join)
原始数据(模拟结构)
| A | B | C |
|---|---|---|
| 12525 | 1FWE23 | 1H654D |
| 14654 | ||
| 24798 | 14654 | S56E82 |
| 65116 | 63546 | 38945 |
| 46456 | 46485 | R68R45 |
| AD545 | ||
| A5D66 | 45346 | QA6683 |
期望输出
| A | B | C |
|---|---|---|
| 12525 | 1FWE23 | 1H654D 14654 |
| 24798 | 14654 | S56E82 |
| 65116 | 63546 | 38945 |
| 46456 | 46485 | R68R45 AD545 |
| A5D66 | 45346 | QA6683 |
报错原因分析
错误的核心原因:
tabula.read_pdf返回的tables[0]并非预期的DataFrame,而是单列Series——这是因为目标PDF表格结构异常,读取时未正确识别多列。- 代码中
groupby使用的列名(如Item、SMR)与实际读取到的列名不匹配,且当输入为Series时,不存在columns属性,直接触发错误。
修正方案
步骤1:优化PDF表格读取
添加stream=True参数让tabula按文本流解析,适配复杂结构的PDF;显式指定列拆分位置,确保正确识别A、B、C三列。
步骤2:统一数据结构与列名
增加类型判断,兼容返回Series的情况,手动重命名列名匹配需求中的A、B、C。
步骤3:空值填充与聚合
用ffill()填充A、B列的空值,再按A、B分组聚合C列内容。
修正后的代码
from tabula import read_pdf import pandas as pd import numpy as np Page_No = 45 # 用stream模式解析,指定列拆分位置(需根据PDF实际列位置调整数值) tables = read_pdf( '/content/210812154731_DECK CRANE - MACGREGOR HAGGLUND - TG 5795-36.5 130-2 - N00024-93-C-2220 - PARTS MANUAL.pdf', pages=Page_No, stream=True, columns=[80, 180, 280] ) # 处理读取结果,确保为DataFrame并匹配列名 if isinstance(tables[0], pd.Series): data_df = tables[0].to_frame() data_df.columns = ['temp'] # 从临时列拆分A、B、C(需根据数据特征调整匹配规则) data_df['A'] = np.where(data_df['temp'].str.len() <= 5, data_df['temp'], np.nan) data_df['B'] = np.where(data_df['temp'].str.match(r'^\w{5,6}$'), data_df['temp'], np.nan) data_df['C'] = np.where(data_df['temp'].str.match(r'^\w{5}$'), data_df['temp'], np.nan) else: data_df = pd.DataFrame(tables[0]) data_df.columns = ['A', 'B', 'C'] # 填充A、B列的空值 data_df[['A', 'B']] = data_df[['A', 'B']].ffill() # 按A、B分组聚合C列 out = data_df.groupby(['A', 'B'], as_index=False)['C'].agg(lambda x: ' '.join([str(i) for i in x if pd.notna(i)])) print(out)
关键说明
columns=[80,180,280]:需根据PDF中A、B、C列的实际位置调整数值,确保tabula正确拆分列。- 列拆分规则:如果PDF中A、B、C列的数据格式有明显特征(如长度、字符类型),可调整
np.where的匹配条件,精准拆分列。 - 空值过滤:聚合时加入空值判断,避免C列空值被拼接成无效内容。
内容的提问来源于stack exchange,提问作者Pravin
相关产品推荐
相关产品推荐

