You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Spacy做文本分词时无法移除特殊字符与多余空白问题求解

问题原因
  • 清洗规则错误删除了缩写的单引号:你定义的字符替换正则中包含了单引号',所有缩写里的单引号都会被替换为空格,it's会被提前处理为it s,无论用什么分词器都会被拆分为两个独立词汇。你所说的NLTK运行正常,大概率是NLTK处理环节没有提前删除单引号,或者用了特殊的合并逻辑。
  • 未处理连续空白:所有匹配到的内容都被替换为单个空格,多轮替换后会产生大量连续空格、首尾空格,spaCy默认分词规则会将连续空格识别为独立的空白token,就会出现你看到的空分词结果。
修正方案

首先调整清洗规则,保留缩写需要的单引号,新增连续空白清理逻辑:

import pandas as pd
import spacy
nlp = spacy.load("en_core_web_sm")

text1 = "[Intro] Well, alright [Chorus] Well, it's 1969, okay? All across the USA It's another year for me and you"
text2 = "[Verse 1] For fifty years they've been married And they can't wait for their fifty-first to roll around"
text3 = "Passion that shouts And red with anger I lost myself Through alleys of mysteries I went up and down Like a demented train"

df = pd.DataFrame({'text':[text1, text2, text3]})

# 修正替换规则:移除正则中的单引号,保留缩写格式
replacer ={
    '\n':' ',
    "\[.*?\]": " ",
    '[!"#%()*+,-./:;<=>?@\[\]^_`{|}~1234567890’”“′‘\\]':" "
}
df['cleanText'] = df['text'].replace(replacer, regex=True)
# 新增:连续空格替换为单个空格,去除首尾空格
df['cleanText'] = df['cleanText'].str.replace('\s+', ' ', regex=True).str.strip()

如果需要将it's这类缩写作为完整的单个分词而不是spaCy默认拆分的it和's,可以给spaCy添加自定义分词规则:

from spacy.symbols import ORTH
# 按需添加需要整词识别的缩写
abbreviation_list = ["it's", "can't", "they've", "don't", "i'm", "we're"]
for abbr in abbreviation_list:
    nlp.tokenizer.add_special_case(abbr, [{ORTH: abbr}])

# 分词测试
df['tokens'] = df['cleanText'].apply(lambda x: [tok.text for tok in nlp(x)])

调整后不会再出现空白分词,缩写的拆分逻辑也符合预期。

内容的提问来源于stack exchange,提问作者PSCM

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.24 16:45:04