DataFrame按Name列左连接触发TypeError问题排查与解决
解决Pandas左连接时的TypeError: type object argument after * must be an iterable, not itertools.imap错误
最近在处理两个Pandas DataFrame的左连接时遇到了一个奇怪的TypeError,报错信息显示type object argument after * must be an iterable, not itertools.imap。折腾了好一阵才找到根源,这里分享一下解决过程。
问题场景
我要对两个DataFrame按Name字段做左连接:
- 第一个DataFrame(145行)包含比赛数据,其中
Name列是通过自定义函数生成的球员姓名 - 第二个DataFrame(1567行)是球员花名册,包含
Name、Team等字段
两个DataFrame的info()输出如下:
第一个DataFrame(比赛数据)
Int64Index: 145 entries, 1 to 162 Data columns (total 13 columns): Quarter 145 non-null float64 Time 145 non-null object Down 145 non-null float64 ToGo 145 non-null float64 Location 145 non-null object Detail 145 non-null object SEA 145 non-null float64 TAM 145 non-null float64 EPB 145 non-null float64 EPA 145 non-null float64 Win% 145 non-null float64 Name 145 non-null object Yards 145 non-null int64 dtypes: float64(8), int64(1), object(4) memory usage: 15.9+ KB
第二个DataFrame(球员花名册)
Int64Index: 1567 entries, 0 to 1566 Data columns (total 4 columns): Name 1567 non-null object Team 1567 non-null object Position 1150 non-null object Age 1567 non-null int64 dtypes: int64(1), object(3) memory usage: 61.2+ KB
执行连接的代码很简单:
df = df.merge(roster, on='Name', how='left')
但触发了以下栈回溯:
Traceback (most recent call last): File "weighted_random_forest.py", line 478, in <module> main() File "weighted_random_forest.py", line 473, in main game = get_third_down_conversion_rate(team, game, files) File "weighted_random_forest.py", line 339, in get_third_down_conversion_rate df = df.merge(roster, on='Name', how='left') File "/usr/local/lib/python2.7/dist-packages/pandas/core/frame.py", line 5370, in merge copy=copy, indicator=indicator, validate=validate) File "/usr/local/lib/python2.7/dist-packages/pandas/core/reshape/merge.py", line 58, in merge return op.get_result() File "/usr/local/lib/python2.7/dist-packages/pandas/core/reshape/merge.py", line 582, in get_result join_index, left_indexer, right_indexer = self._get_join_info() File "/usr/local/lib/python2.7/dist-packages/pandas/core/reshape/merge.py", line 748, in _get_join_info right_indexer) = self._get_join_indexers() File "/usr/local/lib/python2.7/dist-packages/pandas/core/reshape/merge.py", line 727, in _get_join_indexers how=self.how) File "/usr/local/lib/python2.7/dist-packages/pandas/core/reshape/merge.py", line 1050, in _get_join_indexers llab, rlab, shape = map(list, zip(* map(fkeys, left_keys, right_keys))) TypeError: type object argument after * must be an iterable, not itertools.imap
排查过程
一开始我尝试了常规的排查手段:
- 检查两个DataFrame的
Name列数据类型(都是object,没问题) - 调整索引、重置索引
- 删除NaN值
- 甚至修改
itertools的导入方式
但都没有解决问题。
后来通过事后调试器深入查看,发现问题出在左表的Name列中存在空列表元素!
原来我生成Name列的自定义函数name_matcher在匹配失败时返回的是空列表[],而不是空字符串:
原错误函数
def name_matcher(Detail): res = re.findall('[A-Za-z\.\'\-]+ [A-Za-z\'\-]+', Detail) if len(res) != 0: return res[0] else: return [] # 这里返回了空列表,导致Name列混入列表类型元素
Pandas在执行连接时,会尝试对连接键进行哈希和匹配,当键中包含列表这种不可哈希的类型时,就会触发底层的迭代错误,表现为上面的TypeError。
解决方案
只需要修改name_matcher函数,在匹配失败时返回空字符串''而非空列表:
修改后的正确函数
def name_matcher(Detail): res = re.findall('[A-Za-z\.\'\-]+ [A-Za-z\'\-]+', Detail) if len(res) != 0: return res[0] else: return '' # 返回空字符串,保证Name列元素都是字符串类型
修改后重新生成Name列,再执行merge操作就完全正常了。
总结
这个错误的隐蔽之处在于,表面上看是迭代器相关的错误,但根源是连接键列中混入了非字符串(列表)类型的元素。如果遇到类似的Pandas连接错误,一定要检查连接键的元素类型是否统一,避免混入列表、字典等不可哈希的类型。
内容的提问来源于stack exchange,提问作者Brady
相关产品推荐
相关产品推荐

