You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Lambda函数对Pandas数据框的keywords列执行词形还原?

Lemmatize Pandas DataFrame Column with Lambda Function

Hey there! Let's tweak your existing lemmatization code to process the keywords column in your Pandas DataFrame, using a lambda function to apply the logic across every row. Here's the full working solution, plus breakdowns of the key changes:

Full Modified Code

import nltk
# Download required NLTK resources (run once)
nltk.download('wordnet')
nltk.download('averaged_perceptron_tagger')  # Critical for POS tagging
from nltk.stem import WordNetLemmatizer
from nltk.corpus import wordnet
import pandas as pd

# Load your DataFrame
df = pd.read_excel(r'C:\Test2\test.xlsx')

# Initialize lemmatizer
lemmatizer = WordNetLemmatizer()

def get_wordnet_pos(word):
    """Map POS tag to first character lemmatize() accepts"""
    tag = nltk.pos_tag([word])[0][1][0].upper()
    tag_dict = {"J": wordnet.ADJ, "N": wordnet.NOUN, "V": wordnet.VERB, "R": wordnet.ADV}
    return tag_dict.get(tag, wordnet.NOUN)

def lemmatize_single_row(keyword_string):
    # Handle empty/NaN values to avoid errors
    if pd.isna(keyword_string):
        return ""
    # Split string into individual words
    words = keyword_string.split()
    # Lemmatize each word with its correct POS tag
    lemmatized_words = [lemmatizer.lemmatize(word, get_wordnet_pos(word)) for word in words]
    # Join back into a single string
    return ' '.join(lemmatized_words)

# Apply the lemmatization to the entire keywords column using lambda
df['lemmatized_keywords'] = df['keywords'].apply(lambda row: lemmatize_single_row(row))

# Optional: Print a sample to verify results
print(df[['keywords', 'lemmatized_keywords']].head())

Key Changes Explained

  • Added POS tagger download: The nltk.pos_tag() function requires the averaged_perceptron_tagger resource, which was missing in your original code — this prevents runtime errors.
  • Row-specific processing function: lemmatize_single_row() handles splitting the input string, lemmatizing each word, and rejoining the result. It also includes a check for empty/NaN values to keep the code robust.
  • Lambda with apply(): The lambda function acts as a simple wrapper to pass each row's keywords value to our processing function, applying the logic across the entire column efficiently.
  • New output column: We store the lemmatized results in a new column (lemmatized_keywords) so you can compare the original and processed values easily.

This will give you the same accurate lemmatization result as your single-sentence test, but scaled to every row in your DataFrame.

内容的提问来源于stack exchange,提问作者Isa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 08:27:54