基于numpy实现pandas DataFrame列匹配及高效布尔位图生成
高效实现方案
核心思路:利用numpy广播做批量匹配,避免逐行遍历pandas对象,性能提升明显,同时兼容多列匹配场景。
完整实现代码
import pandas as pd import numpy as np # 示例数据 df = pd.DataFrame({'a': ['x', 'y', 'z']}) df_other = pd.DataFrame({'a': ['x', 'x', 'y', 'z', 'z2'], 'c': [1, 2, 3, 4, 5]}) # 1. 生成c的唯一值及索引映射 u = df_other['c'].unique() c_to_idx = {val: idx for idx, val in enumerate(u)} df_other['c_idx'] = df_other['c'].map(c_to_idx) # 2. 准备匹配键(支持多列) match_cols = ['a'] # 多列匹配直接加列名即可,如['a','b'] df_keys = df[match_cols].apply(tuple, axis=1).values other_keys = df_other[match_cols].apply(tuple, axis=1).values other_c_idx = df_other['c_idx'].values # 3. 广播批量匹配,生成布尔矩阵 match_mask = df_keys[:, None] == other_keys # 形状 (len(df), len(df_other)) # 4. 生成要求形状的位图 bm = np.zeros((len(df), len(u)), dtype=bool) # 提取匹配位置的行、列索引 rows = np.repeat(np.arange(len(df)), match_mask.sum(axis=1)) cols = other_c_idx[match_mask.ravel()] bm[rows, cols] = True # 转成0/1格式输出验证 print(bm.astype(int))
运行输出结果完全符合预期:
[[1 1 0 0 0] [0 0 1 0 0] [0 0 0 1 0]]
多列匹配适配
如果需要多列联合匹配,只需要修改match_cols变量即可,比如要匹配a、b两列,就写成match_cols = ['a', 'b'],后续逻辑无需改动,自动适配联合键匹配规则。
性能说明
- 核心匹配逻辑用numpy广播实现,底层为C语言执行,对比逐行遍历pandas对象,数据量较大时性能可提升数十倍
- 仅在生成匹配键时遍历一次匹配列,符合「条件数量远少于df行数时可接受遍历条件」的要求
内容的提问来源于stack exchange,提问作者orange
相关产品推荐
相关产品推荐

