如何为Pandas DataFrame每行的列表列计算众数?
解决DataFrame列表列的每行众数计算问题
你的问题出在df[['B']].mode(axis=1)的用法上:这个方法是对DataFrame的行级元素求众数,但你的B列每个单元格是一个列表(而非单个数值),所以它只会把整个列表当作一个元素返回,自然无法得到你要的结果。
要实现每行列表的众数计算,你需要遍历每个列表,单独计算其众数,这里提供两种可靠的实现方式:
方法1:使用collections.Counter(推荐,无异常风险)
Counter可以直接统计列表中元素的出现次数,再通过max()取出现次数最多的元素:
import pandas as pd from collections import Counter # 构造你的DataFrame(示例) df = pd.DataFrame({ "A": ["Name1", "Name2"], "B": [[3, 5, 6, 6], [1, 1, 1, 4]] }) # 定义众数计算函数 def get_list_mode(lst): count = Counter(lst) # 若有多个众数,返回第一个出现的 return max(count, key=count.get) # 应用到B列生成C列 df['C'] = df['B'].apply(get_list_mode)
方法2:使用statistics.mode(需处理多众数异常)
Python标准库的statistics.mode会直接返回众数,但如果列表存在多个出现次数相同的众数,会抛出StatisticsError,需要额外捕获处理:
import pandas as pd from statistics import mode, StatisticsError def get_list_mode(lst): try: return mode(lst) except StatisticsError: # 多众数时返回第一个出现的高频元素 from collections import Counter count = Counter(lst) max_freq = max(count.values()) return next(k for k, v in count.items() if v == max_freq) df['C'] = df['B'].apply(get_list_mode)
执行后你会得到期望的结果:
A B C 0 Name1 [3, 5, 6, 6] 6 1 Name2 [1, 1, 1, 4] 1
内容的提问来源于stack exchange,提问作者Jasper_97
相关产品推荐
相关产品推荐

