使用Spacy做文本分词时无法移除特殊字符与多余空白问题求解
问题原因
- 清洗规则错误删除了缩写的单引号:你定义的字符替换正则中包含了单引号
',所有缩写里的单引号都会被替换为空格,it's会被提前处理为it s,无论用什么分词器都会被拆分为两个独立词汇。你所说的NLTK运行正常,大概率是NLTK处理环节没有提前删除单引号,或者用了特殊的合并逻辑。 - 未处理连续空白:所有匹配到的内容都被替换为单个空格,多轮替换后会产生大量连续空格、首尾空格,spaCy默认分词规则会将连续空格识别为独立的空白token,就会出现你看到的空分词结果。
修正方案
首先调整清洗规则,保留缩写需要的单引号,新增连续空白清理逻辑:
import pandas as pd import spacy nlp = spacy.load("en_core_web_sm") text1 = "[Intro] Well, alright [Chorus] Well, it's 1969, okay? All across the USA It's another year for me and you" text2 = "[Verse 1] For fifty years they've been married And they can't wait for their fifty-first to roll around" text3 = "Passion that shouts And red with anger I lost myself Through alleys of mysteries I went up and down Like a demented train" df = pd.DataFrame({'text':[text1, text2, text3]}) # 修正替换规则:移除正则中的单引号,保留缩写格式 replacer ={ '\n':' ', "\[.*?\]": " ", '[!"#%()*+,-./:;<=>?@\[\]^_`{|}~1234567890’”“′‘\\]':" " } df['cleanText'] = df['text'].replace(replacer, regex=True) # 新增:连续空格替换为单个空格,去除首尾空格 df['cleanText'] = df['cleanText'].str.replace('\s+', ' ', regex=True).str.strip()
如果需要将it's这类缩写作为完整的单个分词而不是spaCy默认拆分的it和's,可以给spaCy添加自定义分词规则:
from spacy.symbols import ORTH # 按需添加需要整词识别的缩写 abbreviation_list = ["it's", "can't", "they've", "don't", "i'm", "we're"] for abbr in abbreviation_list: nlp.tokenizer.add_special_case(abbr, [{ORTH: abbr}]) # 分词测试 df['tokens'] = df['cleanText'].apply(lambda x: [tok.text for tok in nlp(x)])
调整后不会再出现空白分词,缩写的拆分逻辑也符合预期。
内容的提问来源于stack exchange,提问作者PSCM
相关产品推荐
相关产品推荐

