You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修正DataFrame中词元列表列的词形还原代码以获预期结果?

问题分析与修复方案

错误原因

当前代码的问题在于,lambda row中的row本身就是每个单元格里的词元列表(比如输入的[absolutely, hearts, ...]),但你额外加了一层for y in row,随后对y执行map操作——这里的y是列表里的单个单词,遍历y会把单词拆成单个字符,导致wordnet_lem.lemmatize被逐个应用到字符上,最终输出变成字符列表的嵌套结构。

修复后的代码

直接对整个词元列表应用map即可,无需额外嵌套循环:

from nltk.stem import WordNetLemmatizer
wordnet_lem = WordNetLemmatizer()
df['lemmatized_Tweet'] = df['cleaned_Tweet'].apply(lambda row: list(map(wordnet_lem.lemmatize, row)))

可选优化(提升词形还原精准度)

WordNetLemmatizer默认将单词当作名词处理,若想更精准还原动词、形容词等词性的词,可以给lemmatize方法指定pos参数。比如要把动词形式的calling还原为call,可使用以下代码:

from nltk.corpus import wordnet
from nltk.stem import WordNetLemmatizer
from nltk import pos_tag

def lemmatize_with_pos(token_list):
    wordnet_lem = WordNetLemmatizer()
    # 为每个词标注词性
    tagged_tokens = pos_tag(token_list)
    lemmatized = []
    for token, tag in tagged_tokens:
        # 将NLTK词性映射为WordNet可识别的格式
        pos = wordnet.NOUN
        if tag.startswith('V'):
            pos = wordnet.VERB
        elif tag.startswith('J'):
            pos = wordnet.ADJ
        elif tag.startswith('R'):
            pos = wordnet.ADV
        lemmatized.append(wordnet_lem.lemmatize(token, pos=pos))
    return lemmatized

df['lemmatized_Tweet'] = df['cleaned_Tweet'].apply(lemmatize_with_pos)

内容的提问来源于stack exchange,提问作者Rohan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.07 14:40:49