Python中如何遍历DataFrame按交互ID计算polarity score实现单ID单行输出
实现方案
完整代码
import pandas as pd from nltk.sentiment.vader import SentimentIntensityAnalyzer # 若已提前初始化情感分析器可跳过该行 sid = SentimentIntensityAnalyzer() # 按InteractionId分组聚合,拼接同ID的对话内容 df_grouped = dfo.groupby('InteractionId', as_index=False).agg( Agent=('Agent', 'first'), Transcript=('Transcript', ' '.join) ) # 计算每条完整对话的情感极性分并拆分为独立列 polarity_cols = df_grouped['Transcript'].apply(lambda x: sid.polarity_scores(x)).apply(pd.Series) # 合并原数据与情感分,筛选重命名为你需要的字段 df_result = pd.concat([df_grouped, polarity_cols], axis=1)[ ['InteractionId', 'Agent', 'Transcript', 'pos', 'compound'] ].rename(columns={'pos': 'Positive', 'compound': 'Compound'}) # 可选:导出结果到csv df_result.to_csv("interaction_sentiment.csv", index=False)
逻辑说明
- 分组聚合时,同一个InteractionId对应的Agent完全一致,所以用
first取第一条记录即可,Transcript用空格拼接为完整文本,和你之前遍历CSV拼接全量transcript的逻辑一致 - 调用极性评分接口后用
apply(pd.Series)可以直接把返回的字典格式评分(包含neg、neu、pos、compound四个维度)拆分为独立的数值列 - 最后按需筛选字段重命名,即可得到每个InteractionId对应一行的输出结果
内容的提问来源于stack exchange,提问作者NewInPython
相关产品推荐
相关产品推荐

