使用字典查找函数对比pandas DataFrame两列值并生成匹配结果列
错误原因分析
- 第一种
df.apply写法报错是因为你没指定axis=1,pandas的apply默认参数axis=0,也就是按列遍历,每次传入函数的是整列的Series对象,你尝试从列对象里取x['A']自然会触发索引错误。 - 第二种
assign写法报错是因为assign传入的x是整个DataFrame对象,x['A']拿到的是A列完整的Series,你直接把整个Series作为键去字典里查询,而字典要求键是可哈希的不可变对象,Series是可变对象,因此触发类型错误。
解决方法
方法1:修正现有apply写法
只需要给apply加axis=1参数,指定按行遍历即可:
def do_they_match(A1,A2): if A1 in dictionary and A2 in dictionary and dictionary[A1] == dictionary[A2]: return 1 else: return 0 df['match'] = df.apply(lambda x: do_they_match(x['A'],x['B']), axis=1)
方法2:更高效的向量化实现(推荐,数据量大时性能远高于逐行apply)
用pandas的map方法直接对整列做字典映射,再批量比较:
# 先对两列做字典映射,不存在的键默认返回NaN a_map = df['A'].map(dictionary) b_map = df['B'].map(dictionary) # 比较相等且都不为空时返回1,否则返回0,直接赋值给新列 df['match'] = (a_map.notna() & b_map.notna() & (a_map == b_map)).astype(int)
这种写法不需要自定义函数,也不需要逐行遍历,性能提升非常明显。
内容的提问来源于stack exchange,提问作者Jamie Gorzynski
相关产品推荐
相关产品推荐

