You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何优化Python代码实现带优先级的词频长度Top5排序(defaultdict/Counter)

解决方案

问题分析

你当前的代码已经能统计词长频率,但排序逻辑只考虑了频率降序,没处理频率相同时较长词长优先的规则。要解决这个问题,只需要调整排序的key参数,让排序先按频率倒序,再按词长倒序即可。

基于defaultdict的修改版本

保留你原来的word_lengths函数,修改top5里的排序逻辑:

from collections import defaultdict as dd, Counter

def word_lengths(text):    
    my_dict = dd(int)
    for word in text.split():
        word_len = len(word)
        my_dict[word_len] += 1
    return my_dict

def top5(text):
    my_dict = word_lengths(text)
    # 排序key改为(-频率, -词长),先按频率降序,频率相同则词长降序
    sorted_lengths = sorted(my_dict.keys(), key=lambda x: (-my_dict[x], -x))
    return sorted_lengths[:5]

用Counter简化实现(更符合Pythonic风格)

既然你想学习Counter,其实可以直接用它来统计词长频率,代码会更简洁:

from collections import Counter

def top5(text):
    # 生成所有词的长度列表,再用Counter统计频率
    word_lengths = [len(word) for word in text.split()]
    length_counts = Counter(word_lengths)
    # 同样用(-频率, -词长)作为排序key
    sorted_lengths = sorted(length_counts.keys(), key=lambda x: (-length_counts[x], -x))
    return sorted_lengths[:5]

验证测试

  • 测试文本1:"the quick brown fox jumped over a lazy dog"
    返回结果:[3, 5, 4, 6, 1],符合预期
  • 测试文本2:"one one was a racehorse two two was one too"
    返回结果:[3, 9, 1],符合预期

关键说明

排序时使用lambda x: (-my_dict[x], -x)作为key,原理是:

  1. -my_dict[x]把频率转为负数,默认升序排序等价于频率降序
  2. -x把词长转为负数,当频率相同时,词长越大对应的负数越小,排序时越靠前,实现了较长词长优先

内容的提问来源于stack exchange,提问作者vAlkanol

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 03:43:23