Pandas:计算文本列关键词类别最高概率并生成标签列
解决Pandas文本列关键词类别概率计算及模糊匹配问题
需求说明
- 基于给定三类关键词,为Pandas数据框的
content文本列每行计算各类别的匹配概率:Probability(keyword class) = 该类别匹配关键词的出现次数 / 该行单词总数
- 生成
label列,返回概率最高的类别;概率相同时可返回任意类别,无匹配则返回NaN - 需支持非精确匹配(如
lichies需匹配lichi) - 关键词类别:
fruits = ['mango', 'apple', 'lichi'] animals = ['dog', 'cat', 'cow', 'monkey'] country = ['us', 'ca', 'au', 'br']
原代码问题分析
- 匹配逻辑缺陷:使用
word in fruits的精确匹配,无法处理关键词的变形(如复数lichies) - NaN处理错误:返回字符串
'NaN',而非Pandas标准的缺失值np.nan - 空行边界情况的处理不够严谨
修正后的代码
import re import numpy as np import pandas as pd # 定义关键词类别 fruits = ['mango', 'apple', 'lichi'] animals = ['dog', 'cat', 'cow', 'monkey'] country = ['us', 'ca', 'au', 'br'] def calculate_probability(row): # 提取每行的小写单词列表 words = re.findall(r'\b\w+\b', row['content'].lower()) word_count = len(words) if word_count == 0: return np.nan # 模糊匹配:检查单词是否包含类别中的任意关键词 fruit_matches = sum(any(key in word for key in fruits) for word in words) animal_matches = sum(any(key in word for key in animals) for word in words) country_matches = sum(any(key in word for key in country) for word in words) # 计算各类别概率 probabilities = { 'fruits': fruit_matches / word_count, 'animals': animal_matches / word_count, 'country': country_matches / word_count } # 获取概率最高的类别 max_prob = max(probabilities.values()) if max_prob == 0: return np.nan # 若有多个类别概率相同,返回第一个遇到的 max_label = next(k for k, v in probabilities.items() if v == max_prob) return max_label # 测试用数据框 df = pd.DataFrame({ 'content': [ 'I like mangoes and apples', 'The dog chased the cat', 'US and CA are North American countries', 'Lichies are my favorite fruit', 'Monkey and cow play together', 'No matching keywords here', '' ] }) # 生成label列 df['label'] = df.apply(calculate_probability, axis=1) print(df)
输出结果
content label 0 I like mangoes and apples fruits 1 The dog chased the cat animals 2 US and CA are North American countries country 3 Lichies are my favorite fruit fruits 4 Monkey and cow play together animals 5 No matching keywords here NaN 6 NaN
内容的提问来源于stack exchange,提问作者hxgx_0990
相关产品推荐
相关产品推荐

