Python递归生成pandas DataFrame维度组合路径及对应子表实现方法
现有代码问题分析
现有代码的错误原因:
- 列名和对应值作为两个独立元素存入路径列表,拼接时会多出分隔符,不符合
列名-值的格式要求 - 全局变量回溯逻辑错误,遍历不同分支时残留上一分支的内容,导致路径拼接混乱
- 硬编码层级判断适配性差,更换目标列数量需要修改代码
正确实现方案
使用itertools.product生成全组合即可高效实现需求,无需复杂递归,逻辑清晰不易出错:
import pandas as pd from itertools import product # 示例DataFrame test_list = [['male','pack','lower'], ['male','nonpack','upper'], ['female','pack','upper'], ['male','pack','middle'], ['female','nonpack','middle']] df1 = pd.DataFrame(test_list, columns=['gender', 'subscription', 'ageCategory']) def get_path_df_tuples(df, target_cols): result = [] # 生成每个列的(列名, 唯一值)元组列表 col_value_pairs = [ [(col, val) for val in df[col].unique()] for col in target_cols ] # 遍历所有列值的笛卡尔积组合 for combo in product(*col_value_pairs): # 按要求拼接路径字符串 path = "|".join([f"{col}-{val}" for col, val in combo]) # 筛选对应组合的子DataFrame filter_mask = True for col, val in combo: filter_mask &= (df[col] == val) sub_df = df[filter_mask].reset_index(drop=True) # 追加元组到结果列表 result.append((path, sub_df)) return result # 调用测试 iteration_list = ['gender', 'subscription', 'ageCategory'] list_tuples = get_path_df_tuples(df1, iteration_list) # 打印路径验证输出 for path, _ in list_tuples: print(path)
如果只需要保留原数据中存在匹配记录的组合,可以在追加元组前添加if not sub_df.empty:的判断过滤空结果。
内容的提问来源于stack exchange,提问作者aziz shaw
相关产品推荐
相关产品推荐

