如何在spaCy中为Hashtag设置分词不拆分例外?
解决spaCy不拆分Hashtag的问题
你的代码问题出在add_special_case的用法上:spaCy的add_special_case仅接受具体的字符串作为匹配目标,不支持正则表达式。你传入的r"#\S+"会被当成普通字符串处理,根本匹配不到实际的#hashtag,所以#依然会被拆分成独立token。
下面提供两种可行的解决方案:
方案一:自定义分词器,直接保留完整Hashtag
通过修改spaCy默认分词器的规则,让它把#开头的连续字符识别为一个完整token,不进行拆分:
import re import spacy from spacy.tokenizer import Tokenizer from spacy.util import compile_prefix_regex, compile_infix_regex, compile_suffix_regex class Vectorizer(object): def __init__(self): self.nlp = spacy.load('en_core_web_sm') # 移除默认的#前缀规则(默认会把#单独拆出) prefixes = list(self.nlp.Defaults.prefixes) prefixes.remove(r'#') prefix_re = compile_prefix_regex(prefixes) # 编译保留默认中缀规则(排除可能拆分#的规则) infixes = [x for x in self.nlp.Defaults.infixes if not re.match(r'#', x)] infix_re = compile_infix_regex(infixes) # 自定义分词器,添加hashtag匹配规则 self.nlp.tokenizer = Tokenizer( self.nlp.vocab, prefix_search=prefix_re.search, infix_finditer=infix_re.finditer, suffix_search=compile_suffix_regex(self.nlp.Defaults.suffixes).search, token_match=self._match_hashtag ) def _match_hashtag(self, text): # 匹配以#开头的完整hashtag(直到空格/标点) match = re.fullmatch(r'#\w+', text) return match.group(0) if match else None def tokenize(self, data): lemmatized_tokens = [] for document in data: doc = self.nlp(document) # Hashtag不做词形还原,其他token正常处理 tokens = [ token.text if token.text.startswith('#') else token.lemma_ for token in doc ] lemmatized_tokens.append(tokens) return lemmatized_tokens
测试调用后,输出会是:
[['#hashtag', 'this', 'be', 'a', 'test', ',', 'how', 'do', 'this', 'work']]
方案二:用Matcher合并已拆分的Hashtag
如果不想修改分词器,可以先用默认规则分词,再通过Matcher找到#和后续单词,合并为一个token:
import spacy from spacy.matcher import Matcher class Vectorizer(object): def __init__(self): self.nlp = spacy.load('en_core_web_sm') self.matcher = Matcher(self.nlp.vocab) # 定义匹配规则:# 紧跟一个字母单词 self.matcher.add("HASHTAG", [[{"ORTH": "#"}, {"IS_ALPHA": True}]]) def tokenize(self, data): lemmatized_tokens = [] for document in data: doc = self.nlp(document) # 倒序合并token,避免索引偏移 with doc.retokenize() as retokenizer: for match_id, start, end in reversed(self.matcher(doc)): retokenizer.merge(doc[start:end], attrs={"ORTH": doc[start:end].text}) # 生成结果,Hashtag保留原文本 tokens = [ token.text if token.text.startswith('#') else token.lemma_ for token in doc ] lemmatized_tokens.append(tokens) return lemmatized_tokens
注意事项
Hashtag通常不需要词形还原,所以在生成token列表时要单独判断,直接保留原文本;如果需要对Hashtag内的单词进行还原,可以单独处理(比如token.text.replace(token.text[1:], token.lemma_)),但这一般不符合hashtag的使用场景。
内容的提问来源于stack exchange,提问作者Kilian van Rooijen
相关产品推荐
相关产品推荐

