You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何提升带词性(POS)检查的词形还原(Lemmatisation)速度?

How to Speed Up Lemmatization for Large Text Data in Pandas

Hey there! I get it—waiting for lemmatization to finish on a big DataFrame can be super frustrating. Let's break down why your current code is slow, then walk through actionable fixes to speed things up.

Why Your Current Code Is Slow

Let's spot the bottlenecks first:

  • Reinitializing the lemmatizer every time: You create a new WordNetLemmatizer() inside pre_process, which runs once per row. That's totally unnecessary overhead.
  • Inefficient POS tag handling: You're running pos_tag(words_only) once per row, then looping through indices to fetch tags—this adds redundant index lookups and extra work.
  • Slow synset lookups: The wordnet.synset(...) call for adverbs is relatively expensive, especially if you have lots of adverbs in your data.
  • Row-wise processing limits: Both manual loops and apply process one row at a time, which Pandas isn't optimized for—this drags down speed for large datasets.

Optimizations to Speed Things Up

1. Initialize Resources Only Once

Move the WordNetLemmatizer and wordnet corpus outside the pre_process function so they're created once, not every time you process a row.

2. Streamline POS Tagging & Lemmatization Logic

Instead of looping with indices, iterate directly over words and their POS tags together. This removes redundant lookups and makes the code cleaner.

3. Parallelize Processing with Swifter

Use the swifter library to automatically parallelize the apply operation—it handles optimizing for large datasets without extra hassle.

4. Switch to a Faster Library (spaCy)

For very large datasets, spaCy's lemmatizer is built for speed and handles tokenization, POS tagging, and lemmatization in a single optimized pass, no separate function calls needed.

Optimized Code Examples

Option A: Optimized NLTK Version

from nltk import WordNetLemmatizer, pos_tag
from nltk.corpus import wordnet
import pandas as pd
import swifter  # Install first with pip install swifter

# Initialize resources ONCE outside the function
lem = WordNetLemmatizer()

def pre_process(text):
    words_only = text.lower().split()
    # Get all POS tags in one go
    pos_tags = pos_tag(words_only)
    lemmatized_words = []
    
    for word, tag in pos_tags:
        pos_label = tag[0].lower()
        # Adjust POS label for lemmatizer compatibility
        if pos_label == 'j':
            pos_label = 'a'
        elif pos_label == 'r':
            # Try adverb logic, fall back to default if it fails
            try:
                synset = wordnet.synset(f"{word}.r.1")
                pertainym = synset.lemmas()[0].pertainyms()[0].name()
                lemmatized_words.append(pertainym)
                continue
            except:
                pass
        # Lemmatize based on POS
        if pos_label in ['a', 's', 'v']:
            lemmatized_word = lem.lemmatize(word, pos=pos_label)
        else:
            lemmatized_word = lem.lemmatize(word)
        lemmatized_words.append(lemmatized_word)
    
    return " ".join(lemmatized_words)

# Load your data
df = pd.read_excel('C:/Users/Desktop/TEST.xlsx', sheet_name='Text', engine='openpyxl')

# Parallelize the apply operation
df['Processed Text'] = df['Text'].swifter.apply(pre_process)
clean_text = df['Processed Text'].tolist()

Option B: Faster spaCy Version

SpaCy's pipeline is optimized for speed and reduces redundant steps:

import pandas as pd
import spacy
import swifter

# Load spaCy model (install first with pip install spacy && python -m spacy download en_core_web_sm)
nlp = spacy.load("en_core_web_sm", disable=["parser", "ner"])  # Disable unused components to save speed

def pre_process_spacy(text):
    doc = nlp(text.lower())
    lemmatized_words = []
    for token in doc:
        # Add custom adverb logic here if spaCy's default lemma isn't sufficient
        lemma = token.lemma_
        lemmatized_words.append(lemma)
    return " ".join(lemmatized_words)

# Process the DataFrame
df['Processed Text'] = df['Text'].swifter.apply(pre_process_spacy)
clean_text = df['Processed Text'].tolist()

Key Takeaways

  • Reuse expensive objects: Don't reinitialize lemmatizers or models per row.
  • Avoid row-by-row overhead: Use swifter to parallelize operations and leverage Pandas' strengths.
  • Pick the right tool: SpaCy is often faster than NLTK for large-scale text processing thanks to its optimized pipeline.

内容的提问来源于stack exchange,提问作者123456

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 06:55:11