You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

遍历n行6列DataFrame,统计每行内1~n长度组合的出现频率

解决DataFrame每行元素组合频率统计问题

先看你提供的代码里的几个关键问题:

  • 函数定义语法错误,缺冒号;列表初始化没写等号
  • 循环范围搞错了,应该针对当前行的元素数量(每行6个元素,要生成1到6长度的组合),不是DataFrame的行数
  • 生成组合时用了整个DataFrame而非当前行,逻辑完全错了
  • 函数里的return位置不对,第一次循环就直接返回,根本没收集全所有组合
  • 最后打印的变量名和语法都错了

下面是修正后的可运行代码:

import pandas as pd
import itertools
from collections import Counter

# 创建示例数据(改成6列符合需求)
df = pd.DataFrame([
    [2, 10, 18, 31, 41, 50],
    [12, 27, 28, 39, 42, 55],
    [2, 10, 27, 31, 42, 50]  # 加一行测试重复组合
])

def get_combinations(row):
    all_combinations = []
    # 遍历1到当前行的元素个数(这里是6)的组合长度
    for i in range(1, len(row)+1):
        # 对当前行的元素生成i长度的组合
        combinations = list(itertools.combinations(row, i))
        all_combinations.extend(combinations)
    return all_combinations

# 对每行应用组合生成函数
all_rows_combinations = df.apply(get_combinations, axis=1).tolist()
# 扁平化所有组合列表
all_combinations_flatten = list(itertools.chain.from_iterable(all_rows_combinations))

# 统计每个组合的出现频率
count_combinations = Counter(all_combinations_flatten)

# 打印结果,比如查看几个重复组合的频率
print("部分组合频率:")
print(f"(2,): {count_combinations[(2,)]}")
print(f"(10,): {count_combinations[(10,)]}")
print(f"(2, 10): {count_combinations[(2, 10)]}")
# 也可以遍历所有结果
# for combo, freq in count_combinations.items():
#     print(f"组合{combo}出现{freq}次")

关键修正说明:

  1. 函数修复:补全函数定义的冒号,正确初始化空列表all_combinations = [],循环结束后再返回所有组合
  2. 循环范围:用len(row)获取当前行的元素数量(6),生成1到6长度的组合,符合需求
  3. 组合生成:用itertools.combinations(row, i)对当前行生成组合,而不是整个DataFrame
  4. 数据扁平化:用tolist()把apply的结果转成列表,再用itertools.chain扁平化所有行的组合
  5. 频率统计与打印:正确使用Counter统计,打印时直接访问count_combinations的键值对

运行后就能得到所有行中,从1个元素到6个元素的所有组合的出现次数。

内容的提问来源于stack exchange,提问作者RTrain3k

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 10:23:20