如何将带否定检查的情感词统计函数适配到Pandas DataFrame
适配单篇情感词统计函数到Pandas DataFrame
前提假设
先确认你的单篇文本情感统计函数结构类似如下(如果你的函数逻辑有差异,只需调整核心判断部分即可):
def count_sentiment_words(text, lmdict, negate): # 替换为你实际的分词逻辑(比如jieba分词、NLTK分词等) words = text.strip().split() pos_total = 0 neg_total = 0 for idx, word in enumerate(words): # 处理正面词 if word in lmdict["positive"]: # 取当前词前3个词的范围(避免索引越界) check_window = words[max(0, idx-3):idx] # 窗口内有否定词则计入负面,否则计入正面 if any(neg in check_window for neg in negate): neg_total += 1 else: pos_total += 1 # 处理负面词 elif word in lmdict["negative"]: check_window = words[max(0, idx-3):idx] # 窗口内有否定词则计入正面,否则计入负面 if any(neg in check_window for neg in negate): pos_total += 1 else: neg_total += 1 return pos_total, neg_total
适配到DataFrame的实现
直接使用Pandas的apply方法,对articles列的每一行文本调用上述函数,再将返回的元组拆分为两列(正面词计数、负面词计数):
import pandas as pd # 示例自定义字典和否定词列表 lmdict = { "positive": ["好", "优秀", "满意", "成功"], "negative": ["差", "糟糕", "失败", "不满"] } negate = ["不", "没有", "从未", "并非"] # 示例DataFrame df = pd.DataFrame({ "articles": [ "这个产品非常好,完全超出预期", "我不满意这个结果,它并没有达到要求", "从未见过这么糟糕的服务,简直让人崩溃" ] }) # 应用函数并拆分结果 df[["pos_count", "neg_count"]] = df["articles"].apply( lambda x: pd.Series(count_sentiment_words(x, lmdict, negate)) ) # 查看结果 print(df)
关键说明
- 分词逻辑替换:如果你的文本需要更专业的分词(比如中文的jieba分词),只需修改
count_sentiment_words函数中的words = ...部分即可,例如:import jieba words = list(jieba.cut(text.strip())) - 性能优化:如果你的DataFrame数据量极大(十万级以上行),可以使用
swifter库加速apply操作,只需安装后替换为:import swifter df[["pos_count", "neg_count"]] = df["articles"].swifter.apply( lambda x: pd.Series(count_sentiment_words(x, lmdict, negate)) ) - 结果拆分:使用
pd.Series()将函数返回的元组转为Series,Pandas会自动将其拆分为多列。
内容的提问来源于stack exchange,提问作者dsilva
相关产品推荐
相关产品推荐

