如何提升带词性(POS)检查的词形还原(Lemmatisation)速度?
Hey there! I get it—waiting for lemmatization to finish on a big DataFrame can be super frustrating. Let's break down why your current code is slow, then walk through actionable fixes to speed things up.
Why Your Current Code Is Slow
Let's spot the bottlenecks first:
- Reinitializing the lemmatizer every time: You create a new
WordNetLemmatizer()insidepre_process, which runs once per row. That's totally unnecessary overhead. - Inefficient POS tag handling: You're running
pos_tag(words_only)once per row, then looping through indices to fetch tags—this adds redundant index lookups and extra work. - Slow synset lookups: The
wordnet.synset(...)call for adverbs is relatively expensive, especially if you have lots of adverbs in your data. - Row-wise processing limits: Both manual loops and
applyprocess one row at a time, which Pandas isn't optimized for—this drags down speed for large datasets.
Optimizations to Speed Things Up
1. Initialize Resources Only Once
Move the WordNetLemmatizer and wordnet corpus outside the pre_process function so they're created once, not every time you process a row.
2. Streamline POS Tagging & Lemmatization Logic
Instead of looping with indices, iterate directly over words and their POS tags together. This removes redundant lookups and makes the code cleaner.
3. Parallelize Processing with Swifter
Use the swifter library to automatically parallelize the apply operation—it handles optimizing for large datasets without extra hassle.
4. Switch to a Faster Library (spaCy)
For very large datasets, spaCy's lemmatizer is built for speed and handles tokenization, POS tagging, and lemmatization in a single optimized pass, no separate function calls needed.
Optimized Code Examples
Option A: Optimized NLTK Version
from nltk import WordNetLemmatizer, pos_tag from nltk.corpus import wordnet import pandas as pd import swifter # Install first with pip install swifter # Initialize resources ONCE outside the function lem = WordNetLemmatizer() def pre_process(text): words_only = text.lower().split() # Get all POS tags in one go pos_tags = pos_tag(words_only) lemmatized_words = [] for word, tag in pos_tags: pos_label = tag[0].lower() # Adjust POS label for lemmatizer compatibility if pos_label == 'j': pos_label = 'a' elif pos_label == 'r': # Try adverb logic, fall back to default if it fails try: synset = wordnet.synset(f"{word}.r.1") pertainym = synset.lemmas()[0].pertainyms()[0].name() lemmatized_words.append(pertainym) continue except: pass # Lemmatize based on POS if pos_label in ['a', 's', 'v']: lemmatized_word = lem.lemmatize(word, pos=pos_label) else: lemmatized_word = lem.lemmatize(word) lemmatized_words.append(lemmatized_word) return " ".join(lemmatized_words) # Load your data df = pd.read_excel('C:/Users/Desktop/TEST.xlsx', sheet_name='Text', engine='openpyxl') # Parallelize the apply operation df['Processed Text'] = df['Text'].swifter.apply(pre_process) clean_text = df['Processed Text'].tolist()
Option B: Faster spaCy Version
SpaCy's pipeline is optimized for speed and reduces redundant steps:
import pandas as pd import spacy import swifter # Load spaCy model (install first with pip install spacy && python -m spacy download en_core_web_sm) nlp = spacy.load("en_core_web_sm", disable=["parser", "ner"]) # Disable unused components to save speed def pre_process_spacy(text): doc = nlp(text.lower()) lemmatized_words = [] for token in doc: # Add custom adverb logic here if spaCy's default lemma isn't sufficient lemma = token.lemma_ lemmatized_words.append(lemma) return " ".join(lemmatized_words) # Process the DataFrame df['Processed Text'] = df['Text'].swifter.apply(pre_process_spacy) clean_text = df['Processed Text'].tolist()
Key Takeaways
- Reuse expensive objects: Don't reinitialize lemmatizers or models per row.
- Avoid row-by-row overhead: Use
swifterto parallelize operations and leverage Pandas' strengths. - Pick the right tool: SpaCy is often faster than NLTK for large-scale text processing thanks to its optimized pipeline.
内容的提问来源于stack exchange,提问作者123456

