将不同长度的DataFrame列作为行添加至另一DF时遇错误求助
解决动态列长度下DataFrame行合并的问题
核心问题分析
- 直接用
loc[len(collection)]赋值时,每次从文件读取的目标列长度不一致,导致新行的列数与现有collection的列数不匹配,触发列对齐错误。 - 使用
pd.Series追加时,原temp的索引存在重复值,生成的Series索引不唯一,导致pandas的reindex操作失败,抛出InvalidIndexError。
方法一:自动补全NaN,合并不同长度的行
将目标列转置为单行DataFrame,通过pd.concat合并,pandas会自动对齐所有出现过的列,长度不足的行自动补NaN:
import pandas as pd # 初始化空DataFrame collection = pd.DataFrame() for i in range(len(files)): temp = pd.read_excel( files[i], sheet_name=info["Sheet name"], skiprows=int(info["Rev " + str(rev_num)]["Start"]), nrows=int(info["Rev " + str(rev_num)]["Nrows"]), usecols=info["Rev " + str(rev_num)]["Columns"], header=1 ) # 提取目标列并转置为单行 single_row = temp[info["Data column"]].to_frame().T # 合并到主DataFrame collection = pd.concat([collection, single_row], ignore_index=True)
方法二:将列数据作为列表存入单列
如果不需要将列值拆分为多列,而是把每个文件的目标列数据作为一个整体存储,可将列值转为列表后存入单列:
import pandas as pd # 初始化带指定列的空DataFrame collection = pd.DataFrame(columns=["file_data"]) for i in range(len(files)): temp = pd.read_excel( files[i], sheet_name=info["Sheet name"], skiprows=int(info["Rev " + str(rev_num)]["Start"]), nrows=int(info["Rev " + str(rev_num)]["Nrows"]), usecols=info["Rev " + str(rev_num)]["Columns"], header=1 ) # 将目标列转为列表,添加为一行 collection.loc[len(collection)] = [temp[info["Data column"]].tolist()]
为什么之前的写法会出错?
collection.loc[len(collection)] = temp[...].tolist():赋值的列表长度必须与collection的列数完全一致,一旦后续文件的列长度变化,就会触发列不匹配错误。append(pd.Series(...), ignore_index=True):如果temp的索引存在重复值,生成的Series索引不唯一,pandas在追加时需要reindex,而reindex仅支持唯一索引,因此抛出错误。
内容的提问来源于stack exchange,提问作者jassier
相关产品推荐
相关产品推荐

