spaCy自定义停用词不生效,请求技术支持
Hey there! Let's break down why your custom stop words aren't taking effect—this is a super common pitfall with spaCy, so you’re not alone here.
The Core Issue: Doc Immutability & Order of Operations
Here’s what happened:
- You first ran
df['parsed_transcript'] = df['transcript'].apply(nlp)to convert all your transcripts into spaCyDocobjects. - Then you modified the spaCy vocab to mark your custom words as stop words.
The problem? spaCy’s Doc objects are immutable once created. When you generated those parsed_transcript entries, each token’s is_stop flag was set based on the vocab state at that exact moment. Any changes you make to the vocab later won’t retroactively update those already-existing Docs.
The Fix: Reorder Your Steps
You need to update the vocab with your custom stop words before processing your transcripts into Doc objects. Here’s how to adjust your code:
import spacy import pandas as pd # Step 1: Load spaCy model first nlp = spacy.load("en_core_web_sm") # Step 2: Define your custom stop words (remove spaces like " year "—spaCy matches exact lexemes) my_stop_words = ["thing", "people", "way", "year", "time", "lot", "day"] # Step 3: Mark these words as stop words in the vocab BEFORE processing text for stopword in my_stop_words: # Check if the word exists in the vocab to avoid errors if stopword in nlp.vocab: nlp.vocab[stopword].is_stop = True # Step 4: Now process your transcripts into Doc objects df = pd.read_csv("your_ted_transcripts.csv") # Replace with your actual data load df['parsed_transcript'] = df['transcript'].apply(nlp)
Bonus: If You Can’t Re-Process the Docs (Not Recommended, But Possible)
If for some reason you can’t re-run the apply(nlp) step (e.g., huge dataset), you can manually update the is_stop flag in existing Docs. Note this is less efficient than reprocessing, but here’s how:
def update_stop_words(doc, custom_stops): stop_word_set = set(custom_stops) for token in doc: # Match lowercase to catch capitalized instances too if token.text.lower() in stop_word_set: token.is_stop = True return doc # Apply the update to your existing parsed transcripts df['parsed_transcript'] = df['parsed_transcript'].apply(lambda x: update_stop_words(x, my_stop_words))
Quick Note on Your Stop Word List
Remove entries like " year " (with spaces)—spaCy’s vocab stores individual lexemes, so spaces make this a completely different entry that won’t match any token in your text. Stick to single, clean words.
内容的提问来源于stack exchange,提问作者user3557840

