You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在spaCy中为Hashtag设置分词不拆分例外?

解决spaCy不拆分Hashtag的问题

你的代码问题出在add_special_case的用法上:spaCy的add_special_case仅接受具体的字符串作为匹配目标,不支持正则表达式。你传入的r"#\S+"会被当成普通字符串处理,根本匹配不到实际的#hashtag,所以#依然会被拆分成独立token。

下面提供两种可行的解决方案:

方案一:自定义分词器,直接保留完整Hashtag

通过修改spaCy默认分词器的规则,让它把#开头的连续字符识别为一个完整token,不进行拆分:

import re
import spacy
from spacy.tokenizer import Tokenizer
from spacy.util import compile_prefix_regex, compile_infix_regex, compile_suffix_regex

class Vectorizer(object):
    def __init__(self):
        self.nlp = spacy.load('en_core_web_sm')
        
        # 移除默认的#前缀规则(默认会把#单独拆出)
        prefixes = list(self.nlp.Defaults.prefixes)
        prefixes.remove(r'#')
        prefix_re = compile_prefix_regex(prefixes)
        
        # 编译保留默认中缀规则(排除可能拆分#的规则)
        infixes = [x for x in self.nlp.Defaults.infixes if not re.match(r'#', x)]
        infix_re = compile_infix_regex(infixes)
        
        # 自定义分词器,添加hashtag匹配规则
        self.nlp.tokenizer = Tokenizer(
            self.nlp.vocab,
            prefix_search=prefix_re.search,
            infix_finditer=infix_re.finditer,
            suffix_search=compile_suffix_regex(self.nlp.Defaults.suffixes).search,
            token_match=self._match_hashtag
        )
    
    def _match_hashtag(self, text):
        # 匹配以#开头的完整hashtag(直到空格/标点)
        match = re.fullmatch(r'#\w+', text)
        return match.group(0) if match else None
    
    def tokenize(self, data):
        lemmatized_tokens = []
        for document in data:
            doc = self.nlp(document)
            # Hashtag不做词形还原,其他token正常处理
            tokens = [
                token.text if token.text.startswith('#') else token.lemma_
                for token in doc
            ]
            lemmatized_tokens.append(tokens)
        return lemmatized_tokens

测试调用后,输出会是:

[['#hashtag', 'this', 'be', 'a', 'test', ',', 'how', 'do', 'this', 'work']]

方案二:用Matcher合并已拆分的Hashtag

如果不想修改分词器,可以先用默认规则分词,再通过Matcher找到#和后续单词,合并为一个token:

import spacy
from spacy.matcher import Matcher

class Vectorizer(object):
    def __init__(self):
        self.nlp = spacy.load('en_core_web_sm')
        self.matcher = Matcher(self.nlp.vocab)
        # 定义匹配规则:# 紧跟一个字母单词
        self.matcher.add("HASHTAG", [[{"ORTH": "#"}, {"IS_ALPHA": True}]])
    
    def tokenize(self, data):
        lemmatized_tokens = []
        for document in data:
            doc = self.nlp(document)
            # 倒序合并token,避免索引偏移
            with doc.retokenize() as retokenizer:
                for match_id, start, end in reversed(self.matcher(doc)):
                    retokenizer.merge(doc[start:end], attrs={"ORTH": doc[start:end].text})
            
            # 生成结果,Hashtag保留原文本
            tokens = [
                token.text if token.text.startswith('#') else token.lemma_
                for token in doc
            ]
            lemmatized_tokens.append(tokens)
        return lemmatized_tokens

注意事项

Hashtag通常不需要词形还原,所以在生成token列表时要单独判断,直接保留原文本;如果需要对Hashtag内的单词进行还原,可以单独处理(比如token.text.replace(token.text[1:], token.lemma_)),但这一般不符合hashtag的使用场景。

内容的提问来源于stack exchange,提问作者Kilian van Rooijen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 11:50:31