Python语言检测代码优化与词形还原方案咨询(文本分类预处理)
Got it, let's break down how to fix those two preprocessing bottlenecks you're facing with your Amazon review dataset—slow language detection with TextBlob and suboptimal lemmatization. These are super common pain points when working with large text corpora, so I’ve got some practical, battle-tested solutions for you.
1. Speeding Up Language Detection (Ditching TextBlob's Slow Implementation)
TextBlob’s language detection is notoriously slow for large datasets, mostly because it relies on a relatively heavy underlying library and processes text row-by-row by default. Here are two way faster alternatives:
Option 1: Switch to FastText (Blazing-Fast Local Detection)
FastText is Facebook’s open-source library optimized for text classification tasks, including language detection. It uses a pre-trained local model (no API calls!) and handles batch processing seamlessly, making it 10-15x faster than TextBlob for large datasets.
First, install the library and download the pre-trained language detection model (look for lid.176.bin—it supports 176 languages):
pip install fasttext
Then implement it with your DataFrame:
import fasttext import pandas as pd # Load the pre-trained model (update the path to where you saved lid.176.bin) model = fasttext.load_model('lid.176.bin') # Batch detection function to avoid row-by-row overhead def detect_lang_batch(texts): # Handle missing values first cleaned_texts = [str(t) if pd.notna(t) else "" for t in texts] # Run predictions in batch predictions = model.predict(cleaned_texts) # Extract language code (strip the '__label__' prefix) languages = [pred[0].replace("__label__", "") for pred in predictions[0]] return languages # Apply to your DataFrame df['language'] = detect_lang_batch(df['review_text'].tolist())
Option 2: Parallelize with LangDetect + Swifter
If you prefer a lighter-weight library than FastText, use langdetect (a standalone, faster alternative to TextBlob’s detection) paired with swifter to parallelize processing across CPU cores.
Install dependencies:
pip install langdetect swifter
Code implementation:
from langdetect import detect, LangDetectException import swifter def detect_single_lang(text): try: return detect(text) except LangDetectException: # Handle unrecognizable text return "unknown" # Swifter automatically uses multiprocessing for large DataFrames df['language'] = df['review_text'].swifter.apply(detect_single_lang)
2. Optimizing Lemmatization
Lemmatization can be slow or inaccurate if not implemented properly. Here are the two best optimized approaches:
Option 1: Use spaCy (Best for Speed + Accuracy)
spaCy’s lemmatizer is built for batch processing, includes built-in part-of-speech (POS) tagging (critical for accurate lemmatization), and lets you disable unused components to save memory and speed things up.
Install spaCy and download a lightweight English model:
pip install spacy python -m spacy download en_core_web_sm
Code for batch lemmatization:
import spacy import pandas as pd # Load spaCy model, disable unused components (parser/ner) to speed up processing nlp = spacy.load('en_core_web_sm', disable=['parser', 'ner']) # Batch lemmatization with multiprocessing def lemmatize_batch(texts): lemmatized_texts = [] # Process texts in batches, use all CPU cores with n_process=-1 for doc in nlp.pipe(texts, batch_size=500, n_process=-1): # Keep only alpha tokens, remove stopwords, and lemmatize cleaned_lemmas = [token.lemma_ for token in doc if not token.is_stop and token.is_alpha] lemmatized_texts.append(" ".join(cleaned_lemmas)) return lemmatized_texts # Apply only to English reviews (filter first to save time) english_reviews = df[df['language'] == 'en']['review_text'].tolist() df.loc[df['language'] == 'en', 'lemmatized_text'] = lemmatize_batch(english_reviews)
Option 2: Optimize NLTK's WordNetLemmatizer
If you’re already using NLTK, you can optimize its lemmatizer by adding POS tagging (to improve accuracy) and parallelizing processing with swifter.
First, install NLTK and download required resources:
pip install nltk python -m nltk.download(['wordnet', 'averaged_perceptron_tagger', 'stopwords'])
Optimized code:
from nltk.stem import WordNetLemmatizer from nltk.corpus import wordnet, stopwords from nltk.tag import pos_tag import swifter import pandas as pd lemmatizer = WordNetLemmatizer() stop_words = set(stopwords.words('english')) # Map NLTK POS tags to WordNet POS tags (required for accurate lemmatization) def get_wordnet_pos(tag): if tag.startswith('J'): return wordnet.ADJ elif tag.startswith('V'): return wordnet.VERB elif tag.startswith('N'): return wordnet.NOUN elif tag.startswith('R'): return wordnet.ADV else: default to noun if tag is unrecognized return wordnet.NOUN def lemmatize_single_text(text): if pd.isna(text): return "" # Clean tokens: keep only alpha, remove stopwords tokens = [tok for tok in text.split() if tok.isalpha() and tok.lower() not in stop_words] # Get POS tags for each token tagged_tokens = pos_tag(tokens) # Lemmatize with POS context lemmas = [lemmatizer.lemmatize(tok.lower(), get_wordnet_pos(tag)) for tok, tag in tagged_tokens] return " ".join(lemmas) # Parallelize with swifter df.loc[df['language'] == 'en', 'lemmatized_text'] = df[df['language'] == 'en']['review_text'].swifter.apply(lemmatize_single_text)
内容的提问来源于stack exchange,提问作者Shubham Singh

