pandas如何按指定条件过滤并同时应用函数计算新列
问题描述
现有结构如下的DataFrame:
text, pred score logits No thank you. positive [[0, 0, 1], [1, 0, 2], [1, 0, 0]] [0.01, 0.02, 0.97] They didn't respond me negative [[], [0, 1, 0], [], []] [0.81, 0.10, 0.18]
可使用以下代码生成测试数据:
df = pd.DataFrame({'text':['No thank you', 'They didnt respond me negative'], 'pred':['positive', 'negative'], 'score':['[[0, 0, 1], [1, 0, 2],[1, 0, 0]]', '[[], [0, 1, 0], [], []]'], 'logits':['[0.01, 0.02, 0.97]', '[0.81, 0.10, 0.18]']})
需求说明
- 若
df['pred'] == 'positive',对该行score字段所有子列表的第0位元素求和(空列表按0计算),再乘以该行logits的第2位元素,结果存入新列decision; - 若
df['pred'] == 'negative',对该行score字段所有子列表的第1位元素求和(空列表按0计算),再乘以该行logits的第0位元素,结果存入新列decision。
预期输出如下:
text, pred score logits decision No thank you. positive [[0, 0, 1], [1, 0, 2], [1, 0, 0]] [0.01, 0.02, 0.97] 1.94 They didn't respond me negative [[], [0, 1, 0], [], []] [0.81, 0.10, 0.18] 0.81
问题原因
原有代码无法运行主要有两个核心问题:
score和logits列存储的是字符串格式,无法直接按下标访问元素,需要先转换为Python列表对象- 写法
df['score'][0]是取整个score列的第一行数据,不是逐行取当前行的score值,无法对每行独立计算
正确实现
首先需要用ast.literal_eval把字符串形式的列表转为真实列表,再通过apply逐行计算:
import pandas as pd import ast # 生成测试数据 df = pd.DataFrame({ 'text':['No thank you', 'They didnt respond me negative'], 'pred':['positive', 'negative'], 'score':['[[0, 0, 1], [1, 0, 2],[1, 0, 0]]', '[[], [0, 1, 0], [], []]'], 'logits':['[0.01, 0.02, 0.97]', '[0.81, 0.10, 0.18]'] }) # 将字符串格式的score和logits转为列表对象 df['score'] = df['score'].apply(ast.literal_eval) df['logits'] = df['logits'].apply(ast.literal_eval) # 定义逐行计算逻辑 def get_decision(row): if row['pred'] == 'positive': target_idx = 0 factor = row['logits'][2] else: target_idx = 1 factor = row['logits'][0] total = 0 for sub in row['score']: if sub: total += sub[target_idx] return total * factor # 生成结果列 df['decision'] = df.apply(get_decision, axis=1) print(df)
运行后得到的decision列完全符合预期。
内容的提问来源于stack exchange,提问作者sariii
相关产品推荐
相关产品推荐

