You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何按不同长度预定义子串最长优先拆分字符串?

实现方案

你需要的是最长优先的从左到右子串匹配拆分,这里提供两种可直接运行的实现:

方法1:正则实现(最简写法)

Python的re模块中,正则的交替匹配符|会按模式的书写顺序优先匹配,只要我们把音素列表按长度从大到小排序,就能自动优先匹配更长的子串:

import re

def phonemic_splitter(string):
    phonemes = ['a', 'sh', 's', 'g', 'n', 'c', 'e', 'ch', 'sch']
    # 按音素长度降序排序,保证长串优先匹配
    sorted_phonemes = sorted(phonemes, key=lambda x: len(x), reverse=True)
    # 构建正则模式,re.escape可避免音素含正则特殊字符时报错
    pattern = re.compile('|'.join(re.escape(p) for p in sorted_phonemes))
    return pattern.findall(string)

方法2:手动遍历实现(无需依赖正则)

如果不想引入正则库,可以自己控制匹配逻辑,效率也更高:

def phonemic_splitter(string):
    phonemes = ['a', 'sh', 's', 'g', 'n', 'c', 'e', 'ch', 'sch']
    phoneme_set = set(phonemes)
    max_phoneme_len = max(len(p) for p in phonemes)
    res = []
    current_idx = 0
    str_len = len(string)
    while current_idx < str_len:
        # 从最长可能的长度开始尝试匹配
        for match_len in range(min(max_phoneme_len, str_len - current_idx), 0, -1):
            current_sub = string[current_idx:current_idx+match_len]
            if current_sub in phoneme_set:
                res.append(current_sub)
                current_idx += match_len
                break
    return res

运行验证

两种实现跑你给出的用例都能得到预期结果:

  • phonemic_splitter('case') -> ['c', 'a', 's', 'e']
  • phonemic_splitter('ash') -> ['a', 'sh']
  • phonemic_splitter('change') -> ['ch', 'a', 'n', 'g', 'e']
  • phonemic_splitter('schane') -> ['sch', 'a', 'n', 'e']

内容的提问来源于stack exchange,提问作者Ξένη Γήινος

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.07 06:39:02