如何统计Pandas DataFrame列表列中连续"S"的数量?
问题分析与解决
原代码错误点
- 全局变量未按行重置:
iCount在循环外初始化,遍历所有行时没有为每行重新归零,导致最终所有行都使用最后一次循环结束后的iCount值(第4行计算后的3)。 - 未按行存储结果:没有将每行计算的结果单独保存,而是直接把全局的
iCount赋值给整个iCount列,导致所有行数值完全相同。
修正方案
方法1:修正循环逻辑(逐行计算)
为每行单独初始化计数变量,将每行的计算结果存入列表,最后赋值给DataFrame列:
counts = [] # 遍历col4中的每个列表 for lst in df['col4']: current_count = 0 max_continuous = 0 for item in lst: if item == "S": current_count += 1 # 更新当前行的最大连续S数量 max_continuous = max(max_continuous, current_count) else: # 遇到非S,重置当前连续计数 current_count = 0 counts.append(max_continuous) df['iCount'] = counts
方法2:使用Pandas的apply函数(更简洁)
定义一个计算单列表中连续S最大数量的函数,通过apply逐行处理:
def count_max_continuous_s(lst): current = 0 max_count = 0 for item in lst: if item == "S": current += 1 max_count = max(max_count, current) else: current = 0 return max_count df['iCount'] = df['col4'].apply(count_max_continuous_s)
执行上述任意一种方法后,都能得到期望的结果:
col1 col2 col3 col4 iCount 0 1 6 67 [1, 6, 67] 0 1 S S None [S, S, None] 2 2 3 8 87 [3, 8, 87] 0 3 4 9 79 [4, 9, 79] 0 4 S S S [S, S, S] 3
内容的提问来源于stack exchange,提问作者Giampaolo Levorato
相关产品推荐
相关产品推荐

