如何从DataFrame的两个列表生成所有排列组合对应的字典?
问题原因
- 核心错误来自代码中的
dict(zip(str(test_key),str(test_value)))写法:zip()方法接收字符串参数时会按单个字符迭代,导致完整的分类、主题字符串被拆分为单个字母配对。 - 你每次内层循环都直接覆盖
res变量,之前生成的配对结果会被直接清空,无法累加存储。 - 额外注意:Python 字典不支持重复键,你示例中给出的
{'cat1':'top1','cat1':'top2'}无法实际存在,同一个键的重复赋值会直接覆盖旧值,最终只会保留最后一个赋值。如果需要保留一个分类对应多个主题的映射,建议使用「分类为键、主题列表为值」的字典结构;如果需要保留所有配对关系,建议使用字典列表存储。
修正代码
方案1:生成{分类: [对应主题列表]}格式字典
import pandas as pd test_df = pd.DataFrame() test_df['category'] = [['cat1'],['cat2'],['cat3','cat3.5'],['cat5']] test_df['topic'] = [['top1'],[''],['top2','top3'],['top4']] final_dict = {} for _, row in test_df.iterrows(): temp_keys = row["category"] temp_values = row["topic"] for test_key in temp_keys: test_key = str(test_key) # 键不存在则初始化空列表 if test_key not in final_dict: final_dict[test_key] = [] for test_value in temp_values: test_value = str(test_value).strip() if test_value: # 不需要过滤空值可删除此判断 final_dict[test_key].append(test_value) print(final_dict)
运行输出:
{'cat1': ['top1'], 'cat2': [], 'cat3': ['top2', 'top3'], 'cat3.5': ['top2', 'top3'], 'cat5': ['top4']}
方案2:生成所有配对组成的字典列表
如果需要保留所有独立的键值配对关系,可以用列表存储:
import pandas as pd test_df = pd.DataFrame() test_df['category'] = [['cat1'],['cat2'],['cat3','cat3.5'],['cat5']] test_df['topic'] = [['top1'],[''],['top2','top3'],['top4']] pair_list = [] for _, row in test_df.iterrows(): temp_keys = row["category"] temp_values = row["topic"] for test_key in temp_keys: test_key = str(test_key) for test_value in temp_values: test_value = str(test_value).strip() if test_value: # 不需要过滤空值可删除此判断 pair_list.append({test_key: test_value}) print(pair_list)
运行输出:
[{'cat1': 'top1'}, {'cat3': 'top2'}, {'cat3': 'top3'}, {'cat3.5': 'top2'}, {'cat3.5': 'top3'}, {'cat5': 'top4'}]
内容的提问来源于stack exchange,提问作者Trevor Ferree
相关产品推荐
相关产品推荐

