使用pandas.concat按axis=1拼接却得到axis=0拼接结果的问题
问题描述
尝试横向拼接两个看似索引相同的DataFrame,但结果出现大量NaN,且呈现类似纵向拼接的效果:
代码示例
import pandas as pd import numpy as np dct_l = {'1':'a', '2':'b', '3':'c', '4':'d'} df_l = pd.DataFrame.from_dict(dct_l, orient='index', columns=['Key']) dummy = np.zeros((4,3)) index = np.arange(1,5) columns = ['POW', 'KLA','CSE'] df_e = pd.DataFrame(dummy, index, columns) print(df_l)
输出:
Key 1 a 2 b 3 c 4 d
print(df_e)
输出:
POW KLA CSE 1 0.0 0.0 0.0 2 0.0 0.0 0.0 3 0.0 0.0 0.0 4 0.0 0.0 0.0
执行拼接代码:
pd.concat([df_l, df_e], axis=1)
实际结果:
Key POW KLA CSE 1 a NaN NaN NaN 2 b NaN NaN NaN 3 c NaN NaN NaN 4 d NaN NaN NaN 1 NaN 0.0 0.0 0.0 2 NaN 0.0 0.0 0.0 3 NaN 0.0 0.0 0.0 4 NaN 0.0 0.0 0.0
期望结果:
Key POW KLA CSE 1 a 0.0 0.0 0.0 2 b 0.0 0.0 0.0 3 c 0.0 0.0 0.0 4 d 0.0 0.0 0.0
问题原因
两个DataFrame的索引类型不匹配:
df_l的索引是字符串类型(从字典键'1'、'2'转换而来)df_e的索引是整数类型(np.arange(1,5)生成的整数)
Pandas在按索引横向拼接时,会严格匹配索引的值和类型。虽然表面上索引显示都是1、2、3、4,但类型不同,所以被视为完全不同的索引项,导致每个索引只对应自己DataFrame的列,另一部分列填充NaN,最终呈现出类似纵向拼接的效果。
解决方法
方法1:将df_l的索引转为整数类型
df_l.index = df_l.index.astype(int) pd.concat([df_l, df_e], axis=1)
方法2:创建df_e时使用字符串索引
index = ['1','2','3','4'] # 改用字符串索引 df_e = pd.DataFrame(dummy, index, columns) pd.concat([df_l, df_e], axis=1)
方法3:使用join方法(自动对齐索引,需统一类型)
df_l.index = df_l.index.astype(int) df_l.join(df_e)
内容的提问来源于stack exchange,提问作者RedHand
相关产品推荐
相关产品推荐

