如何用Lambda函数对Pandas数据框的keywords列执行词形还原?
Lemmatize Pandas DataFrame Column with Lambda Function
Hey there! Let's tweak your existing lemmatization code to process the keywords column in your Pandas DataFrame, using a lambda function to apply the logic across every row. Here's the full working solution, plus breakdowns of the key changes:
Full Modified Code
import nltk # Download required NLTK resources (run once) nltk.download('wordnet') nltk.download('averaged_perceptron_tagger') # Critical for POS tagging from nltk.stem import WordNetLemmatizer from nltk.corpus import wordnet import pandas as pd # Load your DataFrame df = pd.read_excel(r'C:\Test2\test.xlsx') # Initialize lemmatizer lemmatizer = WordNetLemmatizer() def get_wordnet_pos(word): """Map POS tag to first character lemmatize() accepts""" tag = nltk.pos_tag([word])[0][1][0].upper() tag_dict = {"J": wordnet.ADJ, "N": wordnet.NOUN, "V": wordnet.VERB, "R": wordnet.ADV} return tag_dict.get(tag, wordnet.NOUN) def lemmatize_single_row(keyword_string): # Handle empty/NaN values to avoid errors if pd.isna(keyword_string): return "" # Split string into individual words words = keyword_string.split() # Lemmatize each word with its correct POS tag lemmatized_words = [lemmatizer.lemmatize(word, get_wordnet_pos(word)) for word in words] # Join back into a single string return ' '.join(lemmatized_words) # Apply the lemmatization to the entire keywords column using lambda df['lemmatized_keywords'] = df['keywords'].apply(lambda row: lemmatize_single_row(row)) # Optional: Print a sample to verify results print(df[['keywords', 'lemmatized_keywords']].head())
Key Changes Explained
- Added POS tagger download: The
nltk.pos_tag()function requires theaveraged_perceptron_taggerresource, which was missing in your original code — this prevents runtime errors. - Row-specific processing function:
lemmatize_single_row()handles splitting the input string, lemmatizing each word, and rejoining the result. It also includes a check for empty/NaN values to keep the code robust. - Lambda with
apply(): The lambda function acts as a simple wrapper to pass each row'skeywordsvalue to our processing function, applying the logic across the entire column efficiently. - New output column: We store the lemmatized results in a new column (
lemmatized_keywords) so you can compare the original and processed values easily.
This will give you the same accurate lemmatization result as your single-sentence test, but scaled to every row in your DataFrame.
内容的提问来源于stack exchange,提问作者Isa
相关产品推荐
相关产品推荐

