如何基于使用频次统一Pandas列中句子内名词的单复数形式
统一处理DataFrame评论中的名词单复数形式
需求说明
给定如下Pandas DataFrame:
import pandas as pd df = pd.DataFrame(data={'ID': [1, 2, 3, 4, 5], 'Comment': ['I love these boots', 'This dog is cute', 'Cool boot', 'This dog like these boots', 'Dogs are cute']})
实际数据包含大量行,评论内容随机。理想目标是:针对句子中的名词,根据其单复数形式的使用频次,统一转换为更常用的形式,期望输出如下:
df = pd.DataFrame(data={'ID': [1, 2, 3, 4, 5], 'Comment': ['I love these boots', 'This dog is cute', 'Cool boots', 'This dog like these boots', 'Dog are cute']})
若无法实现上述理想需求,将所有名词统一转为单数或复数形式也可接受。
实现方案
步骤1:准备NLP工具
使用nltk识别名词并处理单复数,先安装依赖并下载必要资源:
pip install nltk pandas
在Python中下载nltk模型:
import nltk nltk.download('punkt') nltk.download('averaged_perceptron_tagger') nltk.download('wordnet') nltk.download('omw-1.4')
步骤2:提取并统计名词单复数频次
编写函数提取评论中的名词,以单数形式为统一键,统计单复数出现次数:
from nltk.tokenize import word_tokenize from nltk.tag import pos_tag from nltk.stem import WordNetLemmatizer lemmatizer = WordNetLemmatizer() def extract_nouns(text): """提取文本中的单数/复数普通名词""" tokens = word_tokenize(text) tagged = pos_tag(tokens) nouns = [word for word, tag in tagged if tag in ('NN', 'NNS')] return nouns # 提取所有评论中的名词 all_nouns = df['Comment'].apply(extract_nouns).explode().dropna().tolist() # 统计每个名词单复数的出现频次 noun_freq = {} for noun in all_nouns: singular = lemmatizer.lemmatize(noun, pos='n') if singular not in noun_freq: noun_freq[singular] = {'singular': 0, 'plural': 0} if noun == singular: noun_freq[singular]['singular'] += 1 else: noun_freq[singular]['plural'] += 1 # 确定每个名词的标准形式(频次更高的单/复数) standard_nouns = {} for singular, counts in noun_freq.items(): if counts['plural'] > counts['singular']: # 规则复数生成,简单处理;不规则复数可扩展逻辑 standard_nouns[singular] = f"{singular}s" if singular[-1] != 's' else singular else: standard_nouns[singular] = singular # 建立所有名词形式到标准形式的映射 replace_map = {} for singular, standard in standard_nouns.items(): replace_map[singular] = standard plural = f"{singular}s" if singular[-1] != 's' else singular if plural != singular: replace_map[plural] = standard
步骤3:替换评论中的名词
编写替换函数,将每条评论中的名词替换为标准形式:
def replace_nouns(text): tokens = word_tokenize(text) tagged = pos_tag(tokens) new_tokens = [] for word, tag in tagged: if tag in ('NN', 'NNS'): new_tokens.append(replace_map.get(word, word)) else: new_tokens.append(word) return ' '.join(new_tokens) # 应用替换到DataFrame df['Comment'] = df['Comment'].apply(replace_nouns)
简化方案:统一转为单数或复数
若无需根据频次选择,可直接统一转换:
- 统一转为单数:
def to_singular(text): tokens = word_tokenize(text) tagged = pos_tag(tokens) new_tokens = [] for word, tag in tagged: if tag == 'NNS': new_tokens.append(lemmatizer.lemmatize(word, pos='n')) else: new_tokens.append(word) return ' '.join(new_tokens) df['Comment'] = df['Comment'].apply(to_singular)
- 统一转为复数:
def pluralize(word): # 规则复数生成,不规则复数可扩展逻辑 singular = lemmatizer.lemmatize(word, pos='n') return f"{singular}s" if singular[-1] != 's' else singular def to_plural(text): tokens = word_tokenize(text) tagged = pos_tag(tokens) new_tokens = [] for word, tag in tagged: if tag == 'NN': new_tokens.append(pluralize(word)) else: new_tokens.append(word) return ' '.join(new_tokens) df['Comment'] = df['Comment'].apply(to_plural)
注意事项
- 上述代码对规则单复数处理效果较好,针对不规则单复数(如child/children、foot/feet),可结合
wordnet的复杂逻辑或spaCy工具提升准确性。 - 处理大数据量时,建议使用
swifter库加速apply操作:安装pip install swifter后,执行df['Comment'] = df['Comment'].swifter.apply(replace_nouns)。
内容的提问来源于stack exchange,提问作者Mop
相关产品推荐
相关产品推荐

