如何用Python正则表达式修复语料并列省略内容 用于Word2Vec训练
问题原因
你原正则的问题在于重复捕获组( [^ -]+?){1,2}的写法,Python的re模块处理重复捕获组时,只会保留最后一次捕获的内容,因此当连接词和第二个带后缀的单词之间有多个单词时,只会保留最后一个,导致内容丢失。
解决方案
将中间的所有可选单词整体作为一个捕获组,而不是对单个单词做重复捕获,修改后的正则表达式如下:
import re # 正则表达式 pattern = r'(\w+)-\s+(and|to|or)\s+((?:\w+\s+){0,2})(\w+)-([\w-]+)' # 替换规则:把后缀补到第一个前缀后面 result = re.sub(pattern, r'\1-\5 \2 \3\4-\5', text)
规则说明
(\w+):捕获第一个省略后缀的单词前缀(如示例中的seven、pre)\s+(and|to|or)\s+:匹配连接词及前后空格((?:\w+\s+){0,2}):捕获连接词和第二个完整单词之间的0-2个单词,这里用非捕获组(?:)包裹重复逻辑,确保所有中间内容都被完整保留(\w+):捕获第二个单词的前缀(如示例中的eight、post)([\w-]+):捕获后缀部分,兼容带连字符的后缀(如示例中的year-olds、independence)- 替换逻辑直接把后缀补到第一个前缀后,保留中间所有内容
效果验证
输入1:But seeing them playing to seven- and eight-year-olds is beautiful
输出1:But seeing them playing to seven-year-olds and eight-year-olds is beautiful
输入2:The country was in the uproar of pre- and then post-independence civil war but the mood here is most often joyous
输出2:The country was in the uproar of pre-independence and then post-independence civil war but the mood here is most often joyous
内容的提问来源于stack exchange,提问作者abbassix
相关产品推荐
相关产品推荐

