如何高效将numpy数组拼接到pandas DataFrame的每一行
实现方法
方案1:向量化操作(性能最优,适合大数据量)
全程使用numpy向量化运算,无Python层面循环,处理百万行级数据也能保持极高效率:
import pandas as pd import numpy as np # 原始数据 df = pd.DataFrame({'x': range(0,5), 'y' : range(1,6)}) s = np.array(['a', 'b', 'c']) # 核心逻辑 # 按s的长度重复df的每一行 result = df.loc[df.index.repeat(len(s))].copy() # 把s重复对应次数后赋值为新列 result['new_col'] = np.tile(s, len(df)) # 可选:重置为连续索引 result = result.reset_index(drop=True)
方案2:简洁写法(易读性高,适合日常使用)
用pandas内置的assign+explode实现,代码更短,性能满足绝大多数场景需求:
result = df.assign(new_col=[s]*len(df)).explode('new_col', ignore_index=True)
内容的提问来源于stack exchange,提问作者PingPong
相关产品推荐
相关产品推荐

