Pandas处理genre_ids时报ValueError:无法将NA/inf非有限值转为整数
报错原因
- 链式调用时
inplace=True无法作用到原始列:Series.cat.add_categories()会返回一个新的分类序列副本,后续链式调用.fillna(inplace=True)仅修改这个临时副本,原始的train['genre_ids']、test['genre_ids']列仍然存在空值 - 空值传入
genre_id_count函数时,x == 'no_genre_id'判断不成立,执行x.count('|')会返回NaN,最终astype转换int8时触发非有限值转换错误
修复方案
任选以下一种写法即可:
写法1:直接赋值替换原列,去掉inplace参数
# 统计genre_id数量 def genre_id_count(x): if x == 'no_genre_id': return 0 else: return x.count('|') + 1 # 把处理后的序列直接赋值回原列 train['genre_ids'] = train['genre_ids'].cat.add_categories('no_genre_id').fillna('no_genre_id') test['genre_ids'] = test['genre_ids'].cat.add_categories('no_genre_id').fillna('no_genre_id') train['genre_ids_count'] = train['genre_ids'].apply(genre_id_count).astype(np.int8) test['genre_ids_count'] = test['genre_ids'].apply(genre_id_count).astype(np.int8)
写法2:分步骤操作,明确修改原列
# 先给分类列添加新类别 train['genre_ids'].cat.add_categories('no_genre_id', inplace=True) test['genre_ids'].cat.add_categories('no_genre_id', inplace=True) # 再填充空值 train['genre_ids'].fillna('no_genre_id', inplace=True) test['genre_ids'].fillna('no_genre_id', inplace=True) # 后续统计逻辑和原有代码一致 train['genre_ids_count'] = train['genre_ids'].apply(genre_id_count).astype(np.int8) test['genre_ids_count'] = test['genre_ids'].apply(genre_id_count).astype(np.int8)
内容的提问来源于stack exchange,提问作者Heisenberg
相关产品推荐
相关产品推荐

