You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于字典和条件更新DataFrame中的值(取较大值)

问题描述

我有一个DataFrame(df1)和一个字典compare,字典包含df1的部分列名及对应数值。需求是针对df1的['a','c','e']列,将列中每个值与字典对应键的值比较,保留较大值以更新DataFrame。

原始数据

DataFrame df1

>>> df1
    a   b   c   d   e   f
0   4   2   6   2   8   1
1   3   6   7   7   8   5
2   2   1   1   6   8   7
3   1   2   7   3   3   1
4   1   7   2   6   7   6
5   4   8   8   2   2   1

对比字典 compare

compare = {'a':4, 'c':7, 'e':3}

尝试的代码

comp = pd.DataFrame(pd.Series(compare).reindex(df1.columns).fillna(0)).T

df1[df1.columns] = df1.apply(lambda x: np.where(x>comp, x, comp)[0] ,axis=1)

期望输出

>>>df1
    a   b   c   d   e   f
0   4   2   7   2   8   1
1   4   6   7   7   8   5
2   4   1   7   6   8   7
3   4   2   7   3   3   1
4   4   7   7   6   7   6
5   4   8   8   2   3   1

更高效的解决方案

无需构造临时DataFrame或使用apply循环,直接利用pandas的广播特性和np.maximum即可实现需求,代码更简洁且性能更优:

import numpy as np
import pandas as pd

# 指定需要处理的目标列
target_cols = ['a', 'c', 'e']
# 对目标列逐元素取原数据与字典对应值的较大值,更新原DataFrame
df1[target_cols] = np.maximum(df1[target_cols], pd.Series(compare))

方案说明

  • 仅针对目标列操作,避免无意义的全列遍历
  • pd.Series(compare)会自动按列名与df1[target_cols]对齐,实现广播式比较
  • np.maximum是向量化操作,比apply的行循环效率高很多,尤其适合大数据量场景

内容的提问来源于stack exchange,提问作者Yolao_21

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.03 12:10:52