You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

spaCy自定义停用词不生效,请求技术支持

Fixing Custom Stop Words Not Working in spaCy

Hey there! Let's break down why your custom stop words aren't taking effect—this is a super common pitfall with spaCy, so you’re not alone here.

The Core Issue: Doc Immutability & Order of Operations

Here’s what happened:

  • You first ran df['parsed_transcript'] = df['transcript'].apply(nlp) to convert all your transcripts into spaCy Doc objects.
  • Then you modified the spaCy vocab to mark your custom words as stop words.

The problem? spaCy’s Doc objects are immutable once created. When you generated those parsed_transcript entries, each token’s is_stop flag was set based on the vocab state at that exact moment. Any changes you make to the vocab later won’t retroactively update those already-existing Docs.

The Fix: Reorder Your Steps

You need to update the vocab with your custom stop words before processing your transcripts into Doc objects. Here’s how to adjust your code:

import spacy
import pandas as pd

# Step 1: Load spaCy model first
nlp = spacy.load("en_core_web_sm")

# Step 2: Define your custom stop words (remove spaces like " year "—spaCy matches exact lexemes)
my_stop_words = ["thing", "people", "way", "year", "time", "lot", "day"]

# Step 3: Mark these words as stop words in the vocab BEFORE processing text
for stopword in my_stop_words:
    # Check if the word exists in the vocab to avoid errors
    if stopword in nlp.vocab:
        nlp.vocab[stopword].is_stop = True

# Step 4: Now process your transcripts into Doc objects
df = pd.read_csv("your_ted_transcripts.csv") # Replace with your actual data load
df['parsed_transcript'] = df['transcript'].apply(nlp)

If for some reason you can’t re-run the apply(nlp) step (e.g., huge dataset), you can manually update the is_stop flag in existing Docs. Note this is less efficient than reprocessing, but here’s how:

def update_stop_words(doc, custom_stops):
    stop_word_set = set(custom_stops)
    for token in doc:
        # Match lowercase to catch capitalized instances too
        if token.text.lower() in stop_word_set:
            token.is_stop = True
    return doc

# Apply the update to your existing parsed transcripts
df['parsed_transcript'] = df['parsed_transcript'].apply(lambda x: update_stop_words(x, my_stop_words))

Quick Note on Your Stop Word List

Remove entries like " year " (with spaces)—spaCy’s vocab stores individual lexemes, so spaces make this a completely different entry that won’t match any token in your text. Stick to single, clean words.

内容的提问来源于stack exchange,提问作者user3557840

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:49:36