如何阻止spaCy分词器拆分括号包裹的指定内容?
解决spaCy分词器识别方括号标记为独立分词的问题
你添加的特殊规则没生效,核心原因是spaCy的分词逻辑:它会先按默认规则拆分文本,特殊规则只对已经被初步拆分出的完整匹配字符串生效。但你的输入里[intervention]和前后单词连在一起,默认分词器不会把它们拆分,所以特殊规则根本没触发。下面给两种可行的解决办法:
方案1:针对特定标记设置规则
直接修改分词器的前缀、后缀和中缀正则规则,让spaCy能识别[intervention]、[condition]这类标记的边界:
import spacy from spacy.util import compile_prefix_regex, compile_suffix_regex, compile_infix_regex nlp = spacy.load('en_core_web_sm') # 复制默认规则并添加自定义项 prefixes = list(nlp.Defaults.prefixes) suffixes = list(nlp.Defaults.suffixes) infixes = list(nlp.Defaults.infixes) # 添加规则,让标记前后的连接点被识别为分词边界 prefixes.append(r'^\[intervention\]|\[condition\]') suffixes.append(r'\[intervention\]$|\[condition\]$') infixes.append(r'(?<=\w)\[intervention\]|\[condition\](?=\w)') # 重新编译正则并替换分词器规则 nlp.tokenizer.prefix_search = compile_prefix_regex(prefixes).search nlp.tokenizer.suffix_search = compile_suffix_regex(suffixes).search nlp.tokenizer.infix_finditer = compile_infix_regex(infixes).finditer # 测试 text = "A randomized, prospective study of [intervention]endometrial resection[intervention] to prevent [condition]recurrent endometrial polyps[condition]" doc = nlp(text) for token in doc: print(token.text)
运行后,[intervention]和[condition]会被拆成独立分词,不会和前后单词合并。
方案2:批量处理所有方括号包裹内容
如果有大量类似[xxx]的标记,不想逐个指定,用通用正则匹配所有方括号内容:
import spacy from spacy.util import compile_prefix_regex, compile_suffix_regex, compile_infix_regex nlp = spacy.load('en_core_web_sm') prefixes = list(nlp.Defaults.prefixes) suffixes = list(nlp.Defaults.suffixes) infixes = list(nlp.Defaults.infixes) # 通用规则:匹配任何[xxx]格式的内容 prefixes.append(r'^\[[^\]]+\]') suffixes.append(r'\[[^\]]+\]$') infixes.append(r'(?<=\w)\[[^\]]+\]|\[[^\]]+\](?=\w)') # 更新分词器规则 nlp.tokenizer.prefix_search = compile_prefix_regex(prefixes).search nlp.tokenizer.suffix_search = compile_suffix_regex(suffixes).search nlp.tokenizer.infix_finditer = compile_infix_regex(infixes).finditer # 测试 text = "A randomized, prospective study of [intervention]endometrial resection[intervention] to prevent [condition]recurrent endometrial polyps[condition]" doc = nlp(text) for token in doc: print(token.text)
这个方案会把所有[xxx]格式的内容都识别为独立分词,适合批量处理场景。
内容的提问来源于stack exchange,提问作者ignacioct
相关产品推荐
相关产品推荐

