Python中按对应子列表比较嵌套列表并更新计数的方法
嵌套列表的词频计数更新解决方案
明确需求:针对成对的子列表(list_a[0]对应list_b[0],list_a[1]对应list_b[1]),统计list_b中每个字符串在对应list_a子列表里的出现次数,更新list_b的计数,同时统一字符串为小写(示例结果要求)。
高效实现思路
直接遍历检查字符串是否存在会重复扫描list_a子列表,效率较低。更优方式是先为list_a的每个子列表生成词频统计字典,之后通过字典快速查询计数,再更新list_b。
代码实现
list_a = [["the", "ball", "is", "red", "and", "the", "car", "is", "red"], ["the", "boy", "is", "tall", "and", "the", "man", "is", "tall"]] list_b = [[["the", 0], ["ball", 0], ["is", 0], ["red", 0], ["and", 0], ["car", 0]], [["The", 0], ["boy", 0], ["is", 0], ["and", 0], ["man", 0], ["tall", 0]]] # 第一步:为list_a的每个子列表生成词频字典(不区分大小写) word_frequencies = [] for sublist in list_a: freq_dict = {} for word in sublist: lower_word = word.lower() freq_dict[lower_word] = freq_dict.get(lower_word, 0) + 1 word_frequencies.append(freq_dict) # 第二步:遍历list_b,更新每个元素的字符串和计数 for idx, sublist_b in enumerate(list_b): current_freq = word_frequencies[idx] for item in sublist_b: # 统一字符串为小写 lower_word = item[0].lower() item[0] = lower_word # 从词频字典获取计数,不存在则保持0(示例中均存在) item[1] = current_freq.get(lower_word, 0) # 输出验证 print(list_b)
代码说明
- 词频统计:遍历
list_a的每个子列表,将所有词转为小写后统计出现次数,存入字典。后续查询计数的时间复杂度为O(1),大幅提升效率。 - 更新list_b:用
enumerate同时获取list_b子列表的索引和内容,对应到list_a的词频字典。将list_b中的字符串转为小写,并替换计数为词频字典中的值。
执行后list_b将完全符合需求:
[[["the", 2], ["ball", 1], ["is", 2], ["red", 2], ["and", 1], ["car", 1]], [["the", 2], ["boy", 1], ["is", 2], ["and", 1], ["man", 1], ["tall", 2]]]
内容的提问来源于stack exchange,提问作者python_ftw
相关产品推荐
相关产品推荐

