You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas迭代过程中将情感预测字典值添加至DataFrame?

问题描述

我们有一个存储抓取消息的虚拟DataFrame:

import pandas as pd
import numpy as np

temp = pd.DataFrame(np.array([['I am feeling very well',],['It is hard to believe this happened',],
                              ['What is love?',], ['Amazing day today',]]),
                            columns = ['message',])

实际输出的DataFrame内容为:

message            
0    I hate the weather today
1    It is hard to believe this happened
2    What is love
3    Amazing day today

需要迭代每条消息调用情感预测模型提取情感,代码如下:

for i in temp.message:
    x = model.predict(i, 'roberta')

模型返回的结果x是字典格式:

x = {
    "Love" : 0.0931,
    "Hate" : 0.9169,
}

想知道如何在迭代过程中将字典中的所有值添加到原DataFrame中?目前尝试将字典转为DataFrame,但不清楚后续步骤:

for i in temp.message:
    x = model.predict(i, 'roberta')
    y = pd.DataFrame.from_dict(x,orient='index')
    y = y.T
    # what would the next step be?

也考虑过先创建空列再通过消息列左连接,想了解哪种实现方式最优,期望输出如下:

message                                   Love           Hate
0    I hate the weather today                  0.0931         0.9169
1    It is hard to believe this happened       0.444          0.556
...
解决方案

最优实现:列表推导式批量生成结果后合并

直接用列表推导式生成所有消息的预测字典列表,再转为DataFrame与原表横向拼接,这种方法避免了循环中频繁修改DataFrame的低效操作,性能和可读性都最优:

# 生成所有预测结果的字典列表
pred_results = [model.predict(msg, 'roberta') for msg in temp['message']]

# 拼接原表和预测结果表
result_df = pd.concat([temp, pd.DataFrame(pred_results)], axis=1)

循环方式的改进方案(仅作参考,不推荐)

如果一定要用循环,不要每次循环创建小DataFrame再拼接(会产生大量临时对象,性能差),可以采用以下两种方式:

方式1:循环收集字典后批量赋值

pred_list = []
for msg in temp['message']:
    pred_dict = model.predict(msg, 'roberta')
    pred_list.append(pred_dict)

# 将字典列表转为DataFrame后直接赋值给原表的新列
temp[['Love', 'Hate']] = pd.DataFrame(pred_list)

方式2:预先创建空列后逐行赋值

先初始化空列,再通过索引逐行填充:

# 创建空列
temp['Love'] = np.nan
temp['Hate'] = np.nan

# 迭代每条消息的索引和内容,赋值对应预测值
for idx, msg in temp['message'].items():
    pred_dict = model.predict(msg, 'roberta')
    temp.loc[idx, 'Love'] = pred_dict['Love']
    temp.loc[idx, 'Hate'] = pred_dict['Hate']

方案对比

列表推导式的批量处理方式是最优选择,原因如下:

  • 避免了循环中反复修改DataFrame的开销(Pandas的DataFrame为不可变结构,每次拼接都会生成新对象,循环拼接效率极低)
  • 代码简洁直观,维护成本低
  • 批量处理的运行速度远高于逐行循环赋值

内容的提问来源于stack exchange,提问作者DarknessPlusPlus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 09:15:49