You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python正则表达式修复语料并列省略内容 用于Word2Vec训练

问题原因

你原正则的问题在于重复捕获组( [^ -]+?){1,2}的写法,Python的re模块处理重复捕获组时,只会保留最后一次捕获的内容,因此当连接词和第二个带后缀的单词之间有多个单词时,只会保留最后一个,导致内容丢失。

解决方案

将中间的所有可选单词整体作为一个捕获组,而不是对单个单词做重复捕获,修改后的正则表达式如下:

import re
# 正则表达式
pattern = r'(\w+)-\s+(and|to|or)\s+((?:\w+\s+){0,2})(\w+)-([\w-]+)'
# 替换规则:把后缀补到第一个前缀后面
result = re.sub(pattern, r'\1-\5 \2 \3\4-\5', text)

规则说明

  • (\w+):捕获第一个省略后缀的单词前缀(如示例中的seven、pre)
  • \s+(and|to|or)\s+:匹配连接词及前后空格
  • ((?:\w+\s+){0,2}):捕获连接词和第二个完整单词之间的0-2个单词,这里用非捕获组(?:)包裹重复逻辑,确保所有中间内容都被完整保留
  • (\w+):捕获第二个单词的前缀(如示例中的eight、post)
  • ([\w-]+):捕获后缀部分,兼容带连字符的后缀(如示例中的year-olds、independence)
  • 替换逻辑直接把后缀补到第一个前缀后,保留中间所有内容

效果验证

输入1:But seeing them playing to seven- and eight-year-olds is beautiful
输出1:But seeing them playing to seven-year-olds and eight-year-olds is beautiful

输入2:The country was in the uproar of pre- and then post-independence civil war but the mood here is most often joyous
输出2:The country was in the uproar of pre-independence and then post-independence civil war but the mood here is most often joyous

内容的提问来源于stack exchange,提问作者abbassix

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 12:45:01