如何从DataFrame列的类数组元素中提取首个值?
解决DataFrame列提取类数组首个元素的索引越界问题
我的DataFrame某列同时包含单个元素和类数组元素(如numpy数组),想要提取类数组元素的首个值,编写了如下函数:
def get_top_element(x): if isinstance(x, np.ndarray): return x[0] return x df_no_outliers['brand'] = df_no_outliers['brand'].apply(get_top_element)
运行时触发错误:IndexError: index 0 is out of bounds for axis 0 with size 0,该列转为tuple后的结果如截图所示。
问题原因
列中存在空的numpy数组,直接访问索引0会触发越界错误。
解决方法
修改函数,先判断数组是否为空,再进行取值操作:
import numpy as np def get_top_element(x): if isinstance(x, np.ndarray): # 数组非空时取第一个元素,空数组返回None(可替换为你需要的默认值,比如空字符串) return x[0] if x.size > 0 else None return x df_no_outliers['brand'] = df_no_outliers['brand'].apply(get_top_element)
也可以用更简洁的写法,通过np.ravel将多维数组转为一维,再取首个元素:
import numpy as np def get_top_element(x): if isinstance(x, np.ndarray): # 空数组返回None,非空则取第一个元素 return next(iter(np.ravel(x)), None) return x
内容的提问来源于stack exchange,提问作者mustafa00
相关产品推荐
相关产品推荐

