Python 3.10:高效检测字母模式并统计单词单元数
高效统计单词的CV/VC/C/V单元数量(含双字母组合处理)
一、基础实现(无双字母组合)
线性遍历法
这是最直观的线性时间解法,逐个字符检查相邻类型是否不同,相同则单独计数,不同则合并为一个单元。时间复杂度O(n),空间复杂度O(1),适合大多数场景。
def count_units(word: str) -> int: vowels = {'a', 'e', 'i', 'o', 'u', 'y'} word_lower = word.lower() count = 0 i = 0 n = len(word_lower) while i < n: if i + 1 < n: curr_vowel = word_lower[i] in vowels next_vowel = word_lower[i+1] in vowels if curr_vowel != next_vowel: count += 1 i += 2 continue count += 1 i += 1 return count # 验证示例 assert count_units("piece") == 3 assert count_units("queue") == 4 assert count_units("lampshade") == 6
正则表达式法
利用Python的re模块(底层C实现),通过正则优先匹配CV/VC单元,再匹配单个C/V,效率比纯Python遍历更高,尤其适合长单词批量处理。
import re def count_units_regex(word: str) -> int: # 正则优先级:先匹配CV/VC,再匹配单个C/V pattern = r'[bcdfghjklmnpqrstvwxyz][aeiouy]|[aeiouy][bcdfghjklmnpqrstvwxyz]|[bcdfghjklmnpqrstvwxyz]|[aeiouy]' return len(re.findall(pattern, word.lower())) # 验证示例 assert count_units_regex("piece") == 3 assert count_units_regex("queue") == 4 assert count_units_regex("lampshade") == 6
二、进阶实现(支持双字母组合)
对于th/gh这类双字母辅音组合,需要将其视为整体辅音单元,以下两种方案均可实现:
扩展正则法
修改正则模式,优先匹配双字母组合与元音的组合,再处理其他单元,确保双字母组合被当作整体识别。
import re def count_units_with_digraphs(word: str) -> int: # 非捕获组定义双字母辅音组合,避免干扰匹配结果 digraphs = r'(?:th|gh|ch|sh|ph|wh)' # 辅音模式:双字母组合或单个辅音 consonants = rf'{digraphs}|[bcdfghjklmnpqrstvwxyz]' vowels = r'[aeiouy]' # 单元匹配优先级:CV/VC(含双字母)> 单个C/V(含双字母) pattern = rf'{consonants}{vowels}|{vowels}{consonants}|{consonants}|{vowels}' return len(re.findall(pattern, word.lower())) # 验证示例 assert count_units_with_digraphs("thigh") == 2 assert count_units_with_digraphs("piece") == 3 assert count_units_with_digraphs("ghost") == 3
预处理替换法
先将双字母组合替换为单个特殊字符(代表辅音整体),再用基础遍历逻辑处理,规则调整更灵活。
def count_units_with_digraphs_replace(word: str) -> int: vowels = {'a', 'e', 'i', 'o', 'u', 'y'} digraphs = ['th', 'gh', 'ch', 'sh', 'ph', 'wh'] word_lower = word.lower() # 替换双字母组合为特殊辅音标记 for digraph in digraphs: word_lower = word_lower.replace(digraph, '§') # 复用基础遍历逻辑 count = 0 i = 0 n = len(word_lower) while i < n: if i + 1 < n: curr_vowel = word_lower[i] in vowels # 特殊标记视为辅音 next_vowel = word_lower[i+1] in vowels if curr_vowel != next_vowel: count += 1 i += 2 continue count += 1 i += 1 return count # 验证示例 assert count_units_with_digraphs_replace("thigh") == 2 assert count_units_with_digraphs_replace("piece") == 3 assert count_units_with_digraphs_replace("ghost") == 3
内容的提问来源于stack exchange,提问作者Suntooth
相关产品推荐
相关产品推荐

