如何减少Python代码中的for/while循环,提升字符频率统计代码运行速度
代码优化方案
原代码存在两个核心性能瓶颈:
- 最后统计频率的
res.count(i)是O(n)复杂度操作,嵌套在遍历res的循环中整体时间复杂度达到O(n²),文本量越大速度衰减越明显 - while循环实现滑动取子串的写法冗余,Python层的循环次数过多,且额外创建的
res列表会占用不必要的内存
优化后代码
from collections import Counter def main(): text = "Thank you for help me" print(three_letters_count(text)) def three_letters_count(text): def gen_ngrams(): for word in text.lower().split(): if len(word) < 3: yield word continue # 滑动窗口取3字符替代while循环 for idx in range(len(word) - 2): yield word[idx:idx+3] return dict(Counter(gen_ngrams())) if __name__ == '__main__': main()
优化说明
- 用生成器函数产出所有待统计的字符组,不需要额外创建列表存储中间结果,内存占用大幅降低
- 用
range生成滑动窗口下标替代while循环,减少Python层的循环开销 - 调用标准库
Counter做频率统计,底层为C语言实现,比Python层手写的循环统计效率提升10~100倍,文本越长性能增益越明显 - 移除了不必要的
list()转换,split()方法本身返回的就是列表,无额外转换成本
如果追求更精简的写法,也可以用生成器表达式压缩逻辑:
from collections import Counter def three_letters_count(text): ngrams = (word if len(word)<3 else word[idx:idx+3] for word in text.lower().split() for idx in range(max(1, len(word)-2))) return dict(Counter(ngrams))
内容的提问来源于stack exchange,提问作者Fluke
相关产品推荐
相关产品推荐

