You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现字符串字符bigram出现次数统计的技术咨询

统计字符串Bigram出现次数的入门指南

嘿,这事儿不难,我来一步步给你讲清楚怎么用Python搞定这个需求!先明确下:Bigram就是字符串里连续两个字符组成的子串,比如你举的例子"test string",拆分出来的bigram包括te、es、st、t (注意这里t后面是空格)、 s、st……其中st出现了2次,这就是我们要统计的目标。

我给你准备了三种方法,从最基础的手动实现到用内置工具简化,适合入门理解原理:

方法一:手动遍历(适合理解核心逻辑)

这个方法完全靠自己循环生成bigram,用字典记录次数,能帮你彻底搞懂背后的逻辑:

def count_bigrams(s):
    bigram_counts = {}
    # 遍历到倒数第二个字符即可,避免i+1超出字符串长度
    for i in range(len(s) - 1):
        bigram = s[i] + s[i+1]
        # 检查bigram是否已在字典中,存在则计数+1,不存在则初始化为1
        if bigram in bigram_counts:
            bigram_counts[bigram] += 1
        else:
            bigram_counts[bigram] = 1
    return bigram_counts

# 测试你的示例字符串
test_str = "test string"
result = count_bigrams(test_str)
print(result)

运行后输出完全符合你的期望:

{'te': 1, 'es': 1, 'st': 2, 't ': 1, ' s': 1, 'tr': 1, 'ri': 1, 'in': 1, 'ng': 1}

方法二:用collections.defaultdict简化代码

如果你不想写那串if-else判断,可以用Python标准库的defaultdict,它会自动给不存在的键设置默认值,代码更简洁:

from collections import defaultdict

def count_bigrams(s):
    # 初始化字典,默认值为0
    bigram_counts = defaultdict(int)
    for i in range(len(s) - 1):
        bigram = s[i] + s[i+1]
        bigram_counts[bigram] += 1
    # 转成普通字典输出,和第一种方法格式一致
    return dict(bigram_counts)

test_str = "test string"
print(count_bigrams(test_str))

方法三:用collections.Counter(高效简洁的进阶写法)

Python还有个专门用来做计数的工具Counter,能一步完成统计,代码最短:

from collections import Counter

def count_bigrams(s):
    # 用列表推导式生成所有bigram的列表,再传给Counter统计
    bigrams = [s[i:i+2] for i in range(len(s)-1)]
    return dict(Counter(bigrams))

test_str = "test string"
print(count_bigrams(test_str))

额外注意点

  • 空格也是字符!示例里的t 和 s都会被当成有效bigram统计,如果你想忽略空格,可以先对字符串做处理:s = s.replace(" ", "")
  • 大小写敏感:如果字符串有大小写差异,比如"Test String",Te和te会被视为不同的bigram,要是想忽略大小写,可以先转成统一格式:s = s.lower()

这样从基础到进阶的写法,应该能帮你快速掌握bigram统计的方法啦!

内容的提问来源于stack exchange,提问作者Jonas

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 20:42:39