You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python 3.10:高效检测字母模式并统计单词单元数

高效统计单词的CV/VC/C/V单元数量(含双字母组合处理)

一、基础实现(无双字母组合)

线性遍历法

这是最直观的线性时间解法,逐个字符检查相邻类型是否不同,相同则单独计数,不同则合并为一个单元。时间复杂度O(n),空间复杂度O(1),适合大多数场景。

def count_units(word: str) -> int:
    vowels = {'a', 'e', 'i', 'o', 'u', 'y'}
    word_lower = word.lower()
    count = 0
    i = 0
    n = len(word_lower)
    while i < n:
        if i + 1 < n:
            curr_vowel = word_lower[i] in vowels
            next_vowel = word_lower[i+1] in vowels
            if curr_vowel != next_vowel:
                count += 1
                i += 2
                continue
        count += 1
        i += 1
    return count

# 验证示例
assert count_units("piece") == 3
assert count_units("queue") == 4
assert count_units("lampshade") == 6

正则表达式法

利用Python的re模块(底层C实现),通过正则优先匹配CV/VC单元,再匹配单个C/V,效率比纯Python遍历更高,尤其适合长单词批量处理。

import re

def count_units_regex(word: str) -> int:
    # 正则优先级:先匹配CV/VC,再匹配单个C/V
    pattern = r'[bcdfghjklmnpqrstvwxyz][aeiouy]|[aeiouy][bcdfghjklmnpqrstvwxyz]|[bcdfghjklmnpqrstvwxyz]|[aeiouy]'
    return len(re.findall(pattern, word.lower()))

# 验证示例
assert count_units_regex("piece") == 3
assert count_units_regex("queue") == 4
assert count_units_regex("lampshade") == 6

二、进阶实现(支持双字母组合)

对于th/gh这类双字母辅音组合,需要将其视为整体辅音单元,以下两种方案均可实现:

扩展正则法

修改正则模式,优先匹配双字母组合与元音的组合,再处理其他单元,确保双字母组合被当作整体识别。

import re

def count_units_with_digraphs(word: str) -> int:
    # 非捕获组定义双字母辅音组合,避免干扰匹配结果
    digraphs = r'(?:th|gh|ch|sh|ph|wh)'
    # 辅音模式:双字母组合或单个辅音
    consonants = rf'{digraphs}|[bcdfghjklmnpqrstvwxyz]'
    vowels = r'[aeiouy]'
    # 单元匹配优先级:CV/VC(含双字母)> 单个C/V(含双字母)
    pattern = rf'{consonants}{vowels}|{vowels}{consonants}|{consonants}|{vowels}'
    
    return len(re.findall(pattern, word.lower()))

# 验证示例
assert count_units_with_digraphs("thigh") == 2
assert count_units_with_digraphs("piece") == 3
assert count_units_with_digraphs("ghost") == 3

预处理替换法

先将双字母组合替换为单个特殊字符(代表辅音整体),再用基础遍历逻辑处理,规则调整更灵活。

def count_units_with_digraphs_replace(word: str) -> int:
    vowels = {'a', 'e', 'i', 'o', 'u', 'y'}
    digraphs = ['th', 'gh', 'ch', 'sh', 'ph', 'wh']
    word_lower = word.lower()
    
    # 替换双字母组合为特殊辅音标记
    for digraph in digraphs:
        word_lower = word_lower.replace(digraph, '§')
    
    # 复用基础遍历逻辑
    count = 0
    i = 0
    n = len(word_lower)
    while i < n:
        if i + 1 < n:
            curr_vowel = word_lower[i] in vowels
            # 特殊标记视为辅音
            next_vowel = word_lower[i+1] in vowels
            if curr_vowel != next_vowel:
                count += 1
                i += 2
                continue
        count += 1
        i += 1
    return count

# 验证示例
assert count_units_with_digraphs_replace("thigh") == 2
assert count_units_with_digraphs_replace("piece") == 3
assert count_units_with_digraphs_replace("ghost") == 3

内容的提问来源于stack exchange,提问作者Suntooth

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.20 11:44:57