如何用pandas.concat复刻已废弃append方法的原有行为?
问题:用pandas.concat替代append时列数异常
我接手的代码使用了pandas的append方法,触发了官方弃用警告:
The frame.append method is deprecated and will be removed from pandas
in a future version. Use pandas.concat instead.
尝试改用pandas.concat保留原有append行为,但多次尝试均失败:创建一个(0,31)的空DataFrame,用append添加一条空行后结果为(1,31),但各种concat写法都得到(1,32)的结果。
复现代码
import pandas as pd # 创建带列名的空DataFrame obs = pd.DataFrame(columns=['basedatetime_before', 'lat_before', 'lon_before', 'sog_before', 'cog_before', 'heading_before', 'vesselname_before', 'imo_before', 'callsign_before', 'vesseltype_before', 'status_before', 'length_before', 'width_before', 'draft_before', 'cargo_before', 'basedatetime_after', 'lat_after', 'lon_after', 'sog_after', 'cog_after', 'heading_after', 'vesselname_after', 'imo_after', 'callsign_after', 'vesseltype_after', 'status_after', 'length_after', 'width_after', 'draft_after', 'cargo_after']) # 初始化DataFrame desired = pd.Timestamp('2016-03-20 00:05:00+0000', tz='UTC') obs['point'] = desired obs['basedatetime_before'] = pd.to_datetime(obs['basedatetime_before']) obs['basedatetime_after'] = pd.to_datetime(obs['basedatetime_after']) obs.rename(lambda s: s.lower(), axis = 1, inplace = True) # 创建新的"空行"Series new_obs = pd.Series([desired], index=['point']) # 打印初始形状 print("Orig obs.shape", obs.shape) print("New_obs.shape", new_obs.shape) print("--------------------------------------") # 原append写法(正常得到(1,31)) obs1 = obs.append(new_obs, ignore_index=True) # 各种尝试的concat写法(均得到(1,32)) obs2 = pd.concat([obs, new_obs]) obs3 = pd.concat([obs, new_obs], ignore_index=True) obs4 = pd.concat([obs, new_obs.T]) obs5 = pd.concat([obs, new_obs.T], ignore_index=True) obs6 = pd.concat([new_obs, obs]) obs7 = pd.concat([new_obs, obs], ignore_index=True) obs8 = pd.concat([new_obs.T, obs]) obs9 = pd.concat([new_obs.T, obs], ignore_index=True) # 验证append仍正常工作 obs10 = obs.append(new_obs, ignore_index=True) # 打印结果 print("----> obs1.shape",obs1.shape) print("obs2.shape",obs2.shape) print("obs3.shape",obs3.shape) print("obs4.shape",obs4.shape) print("obs5.shape",obs5.shape) print("obs6.shape",obs6.shape) print("obs7.shape",obs7.shape) print("obs8.shape",obs8.shape) print("obs9.shape",obs9.shape) print("----> obs10.shape",obs10.shape)
运行结果
Orig obs.shape (0, 31) New_obs.shape (1,) -------------------------------------- ----> obs1.shape (1, 31) obs2.shape (1, 32) obs3.shape (1, 32) obs4.shape (1, 32) obs5.shape (1, 32) obs6.shape (1, 32) obs7.shape (1, 32) obs8.shape (1, 32) obs9.shape (1, 32) ----> obs10.shape (1, 31)
解决方案
问题根源:直接拼接Series和DataFrame时,concat会将Series的索引当作新列,导致列数增加。而append会自动将Series视为一行,按原DataFrame的列对齐,缺失列填充NaN。
要实现和append完全一致的效果,需先将Series转换为与原DataFrame列对齐的单行DataFrame,再进行拼接:
方法1:先对齐列再转换为DataFrame
# 将new_obs按原DataFrame的列重新索引,缺失列自动填充NaN aligned_new_obs = new_obs.reindex(obs.columns) # 转换为单行DataFrame后拼接,设置ignore_index=True(和原append参数对应) obs_concat = pd.concat([obs, pd.DataFrame([aligned_new_obs])], ignore_index=True) print("obs_concat.shape", obs_concat.shape) # 输出 (1, 31)
方法2:直接指定列名创建DataFrame
# 直接用原DataFrame的列名创建单行DataFrame,缺失列自动填充NaN obs_concat = pd.concat([obs, pd.DataFrame([new_obs], columns=obs.columns)], ignore_index=True) print("obs_concat.shape", obs_concat.shape) # 输出 (1, 31)
两种写法都能得到和append完全一致的结果,保证列数为31。
内容的提问来源于stack exchange,提问作者user1245262
相关产品推荐
相关产品推荐

