Pandas中concat与append的结果结构是否存在差异?
问题
我查过Stack Overflow相关帖子但没解决问题:用append和concat各添加4条相同行,结果视觉上不一致;提取数据列时也有差异。想确认是仅显示不同还是用法有误?我需要的是append生成的结果。可复现代码如下:
import pandas as pd import numpy as np from datetime import datetime from string import ascii_lowercase as al # artificial dataframe np.random.seed(365) rows = 15 cols = 2 data = np.random.randint(0, 10, size=(rows, cols)) index = pd.bdate_range(datetime.today(), freq='d', periods=rows) dfdata = pd.DataFrame(data=data, index=index, columns=list(al[:cols])) # dfdata = pd.read_csv("data.csv") dfdatanew = pd.DataFrame() frames = [dfdata.iloc[2],dfdata.iloc[2],dfdata.iloc[2],dfdata.iloc[2]] dfdatanew = dfdatanew.append(dfdata.iloc[2]) dfdatanew = dfdatanew.append(dfdata.iloc[2]) dfdatanew = dfdatanew.append(dfdata.iloc[2]) dfdatanew = dfdatanew.append(dfdata.iloc[2]) result = pd.concat(frames,axis=0,join='outer') # compare print(result) print(dfdatanew)
原因分析
两者的视觉差异本质是索引结构不同,并非数据内容有区别:
append处理单个Series(dfdata.iloc[2]是Series)时,会自动把Series转为单行DataFrame,并生成递增的整数索引。concat直接拼接多个Series时,会保留每个Series原有的日期索引,因为4条数据的索引完全重复,输出时会显示成分层缩进的格式,看起来和append的结果不一样。
解决方法
要让concat生成和append一致的结果,有两种简单方式:
方式1:拼接前将Series转为单行DataFrame
把dfdata.iloc[2]改成dfdata.iloc[2:3](切片得到单行DataFrame),同时添加ignore_index=True参数重置索引:
frames = [dfdata.iloc[2:3], dfdata.iloc[2:3], dfdata.iloc[2:3], dfdata.iloc[2:3]] result = pd.concat(frames, axis=0, join='outer', ignore_index=True)
方式2:拼接后重置索引
直接对concat的结果调用reset_index(drop=True),丢弃原有重复索引并生成新的整数索引:
frames = [dfdata.iloc[2],dfdata.iloc[2],dfdata.iloc[2],dfdata.iloc[2]] result = pd.concat(frames, axis=0, join='outer').reset_index(drop=True)
两种方式得到的result和append生成的dfdatanew结构、内容完全一致,提取列的结果也会相同。
内容的提问来源于stack exchange,提问作者Jesse Feng
相关产品推荐
相关产品推荐

