You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于numpy实现pandas DataFrame列匹配及高效布尔位图生成

高效实现方案

核心思路:利用numpy广播做批量匹配,避免逐行遍历pandas对象,性能提升明显,同时兼容多列匹配场景。

完整实现代码

import pandas as pd
import numpy as np

# 示例数据
df = pd.DataFrame({'a': ['x', 'y', 'z']})
df_other = pd.DataFrame({'a': ['x', 'x', 'y', 'z', 'z2'], 'c': [1, 2, 3, 4, 5]})

# 1. 生成c的唯一值及索引映射
u = df_other['c'].unique()
c_to_idx = {val: idx for idx, val in enumerate(u)}
df_other['c_idx'] = df_other['c'].map(c_to_idx)

# 2. 准备匹配键(支持多列)
match_cols = ['a']  # 多列匹配直接加列名即可,如['a','b']
df_keys = df[match_cols].apply(tuple, axis=1).values
other_keys = df_other[match_cols].apply(tuple, axis=1).values
other_c_idx = df_other['c_idx'].values

# 3. 广播批量匹配,生成布尔矩阵
match_mask = df_keys[:, None] == other_keys  # 形状 (len(df), len(df_other))

# 4. 生成要求形状的位图
bm = np.zeros((len(df), len(u)), dtype=bool)
# 提取匹配位置的行、列索引
rows = np.repeat(np.arange(len(df)), match_mask.sum(axis=1))
cols = other_c_idx[match_mask.ravel()]
bm[rows, cols] = True

# 转成0/1格式输出验证
print(bm.astype(int))

运行输出结果完全符合预期:

[[1 1 0 0 0]
 [0 0 1 0 0]
 [0 0 0 1 0]]

多列匹配适配

如果需要多列联合匹配,只需要修改match_cols变量即可,比如要匹配a、b两列,就写成match_cols = ['a', 'b'],后续逻辑无需改动,自动适配联合键匹配规则。

性能说明

  • 核心匹配逻辑用numpy广播实现,底层为C语言执行,对比逐行遍历pandas对象,数据量较大时性能可提升数十倍
  • 仅在生成匹配键时遍历一次匹配列,符合「条件数量远少于df行数时可接受遍历条件」的要求

内容的提问来源于stack exchange,提问作者orange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.25 03:06:04