You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于使用频次统一Pandas列中句子内名词的单复数形式

统一处理DataFrame评论中的名词单复数形式

需求说明

给定如下Pandas DataFrame:

import pandas as pd
df = pd.DataFrame(data={'ID': [1, 2, 3, 4, 5], 'Comment': ['I love these boots', 'This dog is cute', 'Cool boot', 'This dog like these boots', 'Dogs are cute']})

实际数据包含大量行,评论内容随机。理想目标是:针对句子中的名词,根据其单复数形式的使用频次,统一转换为更常用的形式,期望输出如下:

df = pd.DataFrame(data={'ID': [1, 2, 3, 4, 5], 'Comment': ['I love these boots', 'This dog is cute', 'Cool boots', 'This dog like these boots', 'Dog are cute']})

若无法实现上述理想需求,将所有名词统一转为单数或复数形式也可接受。

实现方案

步骤1:准备NLP工具

使用nltk识别名词并处理单复数,先安装依赖并下载必要资源:

pip install nltk pandas

在Python中下载nltk模型:

import nltk
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')
nltk.download('wordnet')
nltk.download('omw-1.4')

步骤2:提取并统计名词单复数频次

编写函数提取评论中的名词,以单数形式为统一键,统计单复数出现次数:

from nltk.tokenize import word_tokenize
from nltk.tag import pos_tag
from nltk.stem import WordNetLemmatizer

lemmatizer = WordNetLemmatizer()

def extract_nouns(text):
    """提取文本中的单数/复数普通名词"""
    tokens = word_tokenize(text)
    tagged = pos_tag(tokens)
    nouns = [word for word, tag in tagged if tag in ('NN', 'NNS')]
    return nouns

# 提取所有评论中的名词
all_nouns = df['Comment'].apply(extract_nouns).explode().dropna().tolist()

# 统计每个名词单复数的出现频次
noun_freq = {}
for noun in all_nouns:
    singular = lemmatizer.lemmatize(noun, pos='n')
    if singular not in noun_freq:
        noun_freq[singular] = {'singular': 0, 'plural': 0}
    if noun == singular:
        noun_freq[singular]['singular'] += 1
    else:
        noun_freq[singular]['plural'] += 1

# 确定每个名词的标准形式(频次更高的单/复数)
standard_nouns = {}
for singular, counts in noun_freq.items():
    if counts['plural'] > counts['singular']:
        # 规则复数生成,简单处理;不规则复数可扩展逻辑
        standard_nouns[singular] = f"{singular}s" if singular[-1] != 's' else singular
    else:
        standard_nouns[singular] = singular

# 建立所有名词形式到标准形式的映射
replace_map = {}
for singular, standard in standard_nouns.items():
    replace_map[singular] = standard
    plural = f"{singular}s" if singular[-1] != 's' else singular
    if plural != singular:
        replace_map[plural] = standard

步骤3:替换评论中的名词

编写替换函数,将每条评论中的名词替换为标准形式:

def replace_nouns(text):
    tokens = word_tokenize(text)
    tagged = pos_tag(tokens)
    new_tokens = []
    for word, tag in tagged:
        if tag in ('NN', 'NNS'):
            new_tokens.append(replace_map.get(word, word))
        else:
            new_tokens.append(word)
    return ' '.join(new_tokens)

# 应用替换到DataFrame
df['Comment'] = df['Comment'].apply(replace_nouns)

简化方案:统一转为单数或复数

若无需根据频次选择,可直接统一转换:

  • 统一转为单数:
def to_singular(text):
    tokens = word_tokenize(text)
    tagged = pos_tag(tokens)
    new_tokens = []
    for word, tag in tagged:
        if tag == 'NNS':
            new_tokens.append(lemmatizer.lemmatize(word, pos='n'))
        else:
            new_tokens.append(word)
    return ' '.join(new_tokens)

df['Comment'] = df['Comment'].apply(to_singular)
  • 统一转为复数:
def pluralize(word):
    # 规则复数生成,不规则复数可扩展逻辑
    singular = lemmatizer.lemmatize(word, pos='n')
    return f"{singular}s" if singular[-1] != 's' else singular

def to_plural(text):
    tokens = word_tokenize(text)
    tagged = pos_tag(tokens)
    new_tokens = []
    for word, tag in tagged:
        if tag == 'NN':
            new_tokens.append(pluralize(word))
        else:
            new_tokens.append(word)
    return ' '.join(new_tokens)

df['Comment'] = df['Comment'].apply(to_plural)

注意事项

  • 上述代码对规则单复数处理效果较好,针对不规则单复数(如child/children、foot/feet),可结合wordnet的复杂逻辑或spaCy工具提升准确性。
  • 处理大数据量时,建议使用swifter库加速apply操作:安装pip install swifter后,执行df['Comment'] = df['Comment'].swifter.apply(replace_nouns)。

内容的提问来源于stack exchange,提问作者Mop

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 09:30:32