Python函数分组逻辑异常:匹配子列表元素生成字典不符合预期
修正group_by_category函数的逻辑问题
需求说明
编写group_by_category函数,检查colors子列表的元素是否存在于food的子列表文本中,基于categ_cols中的颜色首字母生成字典。
现有代码
def group_by_category(categs_names, categs_subs, text_subs): temp_dict = {name: [] for name in categs_names} for categ_index, category_list in enumerate(categs_subs): for substring in category_list: for sublist in text_subs: for text in sublist: if substring in text: temp_dict[categs_names[categ_index]].append(text) else: temp_dict[categs_names[categ_index]].append('None') return temp_dict categs_cols = ['y', 'r', 'g'] food = [['banana is yellow', 'apple is red', 'pear is green' ] , ['lettuce is green' ,'pasta is yellowish','sugar is brown'] ] colors = [[ 'yellow', 'yellowish' ], ['red'], ['green']] grouped_cols = group_by_category(categs_cols, colors, food) print(grouped_cols)
预期与实际结果
预期结果
{'y': ['banana is yellow', 'pasta is yellowish'], 'r': ['None', 'apple is red'], 'g': ['pear is green', 'lettuce is green']}
实际运行结果
{'y': ['banana is yellow', 'None', 'None', 'None', 'pasta is yellowish', 'None', 'None', 'None', 'None', 'None', 'pasta is yellowish', 'None'], 'r': ['None', 'apple is red', 'None', 'None', 'None', 'None'], 'g': ['None', 'None', 'pear is green', 'lettuce is green', 'None', 'None']}
问题分析
原代码的核心问题是循环逻辑混乱:
- 多层嵌套循环导致同一文本被重复处理:每遍历一个颜色子串,就会对所有文本执行一次“匹配则添加文本,不匹配则添加'None'”的操作,产生大量冗余的
None和重复的匹配项。 - 没有实现“每个文本仅针对对应类别判断一次”的逻辑,而是无差别重复遍历。
修正后的代码
def group_by_category(categs_names, categs_subs, text_subs): # 建立类别与对应颜色子串的映射 categ_sub_map = dict(zip(categs_names, categs_subs)) result = {name: [] for name in categs_names} # 遍历每个food子列表,为每个类别生成对应结果 for sublist in text_subs: for categ, substrings in categ_sub_map.items(): matched_text = None # 在当前子列表中查找第一个匹配的文本 for text in sublist: if any(sub in text for sub in substrings): matched_text = text break # 加入结果,无匹配则填'None' result[categ].append(matched_text if matched_text else 'None') return result categs_cols = ['y', 'r', 'g'] food = [['banana is yellow', 'apple is red', 'pear is green' ] , ['lettuce is green' ,'pasta is yellowish','sugar is brown'] ] colors = [[ 'yellow', 'yellowish' ], ['red'], ['green']] grouped_cols = group_by_category(categs_cols, colors, food) print(grouped_cols)
运行结果
{'y': ['banana is yellow', 'pasta is yellowish'], 'r': ['apple is red', 'None'], 'g': ['pear is green', 'lettuce is green']}
注:此结果与用户预期的r项存在差异,推测是用户预期结果的笔误。如果需要r项为['None', 'apple is red'],可调整子列表的遍历顺序或匹配规则。
内容的提问来源于stack exchange,提问作者Gogo
相关产品推荐
相关产品推荐

