Pandas逐行对比name与new-name列 按条件筛选输出指定列
pandas 逐行多条件筛选DataFrame指定字段实现方案
你之前使用的df[person-name].equals(df[new-person])作用是判断两个列的整体内容是否完全一致,无法支持逐行自定义多条件匹配,直接使用pandas原生布尔索引即可实现需求,不需要逐行循环。
核心实现代码
# 构造逐行筛选条件:同时满足name列值为sidney、new-name列值为SC # 注意每个单条件必须加圆括号,避免运算符优先级报错 filter_condition = (df['name'] == 'sidney') & (df['new-name'] == 'SC') # 筛选符合条件的行,提取指定的4个目标字段 target_result = df.loc[filter_condition, ['identifier', 'person-name', 'new-identifier', 'new-person']]
示例验证
用你给出的样例数据构造测试DataFrame:
import pandas as pd # 构造测试数据集 df = pd.DataFrame([ { 'identifier': ('hockey', 'player'), 'person-name': 'sidney crosby', 'type': 'player', 'name': 'sidney', 'new-identifier': ('pittsburg', 'player'), 'new-person': 'crosby-sidney', 'new-type': 'player', 'new-name': 'SC' }, { 'identifier': ('basketball', 'player'), 'person-name': 'lebron james', 'type': 'player', 'name': 'lebron', 'new-identifier': ('lakers', 'player'), 'new-person': 'james-lebron', 'new-type': 'player', 'new-name': 'LJ' } ])
运行上述筛选代码后,输出结果完全匹配预期:
- identifier:
('hockey', 'player') - person-name:
sidney crosby - new-identifier:
('pittsburg', 'player') - new-person:
crosby-sidney
注意事项
- 列名包含横杠等特殊字符时,必须用字符串引号包裹列名,直接写
df[person-name]会被识别为变量运算触发报错- 多条件组合时,
&代表逻辑与(所有条件同时满足)、|代表逻辑或(满足任意条件即可)、~代表逻辑非(条件取反),每个独立判断条件必须用圆括号包裹- 布尔索引是pandas向量化实现,性能远高于
apply逐行遍历、for循环遍历的写法,数据量越大优势越明显
内容的提问来源于stack exchange,提问作者gcoder
相关产品推荐
相关产品推荐

