You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现统计文件中以www开头的唯一字符串出现次数

嘿,刚入门Python是吧?这个统计以www开头字符串唯一出现次数的需求其实挺典型的,我给你捋个清晰的实现思路,附带可直接用的代码,保证你能看懂~

核心思路拆解

其实就四步,很简单:

  • 读取文件内容,把内容拆分成单个字符串(按空白字符分割,包括空格、换行、制表符)
  • 筛选出所有以www开头的字符串
  • 用Python自带的工具统计每个字符串的出现次数
  • 按你要的格式输出结果

基础实现代码(适合小文件)

这个版本适合文件不大的情况,代码简洁易懂:

from collections import Counter

def count_www_strings(file_path):
    # 用with语句打开文件,会自动关闭,避免资源泄漏
    with open(file_path, 'r', encoding='utf-8') as f:
        # 读取全部内容,按空白字符分割成字符串列表
        all_words = f.read().split()
    
    # 用列表推导式筛选出以www开头的字符串
    www_words = [word for word in all_words if word.startswith('www')]
    
    # 用Counter统计每个字符串的出现次数,这是Python标准库的工具,超方便
    word_counts = Counter(www_words)
    
    # 循环输出结果,用f-string格式化字符串
    for word, count in word_counts.items():
        print(f"{word} - {count}")

# 把这里的'your_file.txt'替换成你实际的文件路径,比如'test.txt'
count_www_strings('your_file.txt')

进阶版(适合大文件)

如果你的文件特别大(比如几百MB甚至几GB),一次性把所有内容读入内存可能会卡顿,那就用逐行处理的版本,内存占用会小很多:

from collections import Counter

def count_www_strings_large_file(file_path):
    # 先初始化一个Counter对象用来统计
    word_counts = Counter()
    
    with open(file_path, 'r', encoding='utf-8') as f:
        # 逐行读取文件,处理完一行就释放一行的内存
        for line in f:
            # 分割当前行的字符串
            words = line.split()
            # 筛选出当前行里以www开头的字符串
            www_words = [word for word in words if word.startswith('www')]
            # 更新统计结果
            word_counts.update(www_words)
    
    # 输出最终统计结果
    for word, count in word_counts.items():
        print(f"{word} - {count}")

count_www_strings_large_file('large_file.txt')

测试示例

比如你给的示例内容:

www_1.youtube.com www_1.youtube.com www_3.google.com www_1.youtube.com

把这段内容保存到test.txt里,运行上面的代码,就会得到你想要的输出:
www_1.youtube.com - 3
www_3.google.com - 1

内容的提问来源于stack exchange,提问作者DeepG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 08:16:46