You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何实现类似字典update方法的DataFrame更新效果?

实现类似字典update的DataFrame更新效果

Python 字典的update方法可以基于另一个字典更新已有键值,同时添加新的键值对:

d1 = {"asd": 0, "lol": 1}
d2 = {"lol": 2, "foo": 3}
d1.update(d2)
assert d1 == {'asd': 0, 'lol': 2, 'foo': 3}

想要在 Pandas DataFrame 中实现相同逻辑:更新重叠索引的列值,同时添加新行。现有数据如下:

import pandas as pd

df1 = pd.DataFrame(0, index=range(3), columns=["a"])
# df1 输出:
#    a
# 0  0
# 1  0
# 2  0

df2 = pd.DataFrame(1, index=range(2, 4), columns=["a"])
# df2 输出:
#    a
# 2  1
# 3  1

直接调用df1.update(df2)无法得到预期结果:

df1.update(df2)
# 更新后的 df1:
#      a
# 0  0.0
# 1  0.0
# 2  1.0

不仅整数被自动转为浮点数,df2 独有的索引3行也未被添加。预期输出应为:

a
0   0
1   0
2   1
3   1

解决方案

方法1:使用combine_first

combine_first会用 df2 的值覆盖 df1 重叠索引的内容,同时补充 df1 没有的行,最后通过astype(int)恢复整数类型:

result = df2.combine_first(df1).astype(int)
print(result)

输出完全符合预期:

a
0  0
1  0
2  1
3  1

方法2:合并后保留最新值

先将两个 DataFrame 按行合并,再按索引分组并保留每组最后一行(确保 df2 的值优先级更高),最后排序索引:

result = pd.concat([df1, df2]).groupby(level=0).last().astype(int).sort_index()
print(result)

同样能得到预期结果。


关键说明

  • Pandas 的df.update()设计逻辑是仅更新已有索引和列的对应值,不会新增行或列,这是它和字典update的核心差异。
  • 整数转浮点数是因为更新操作会引入临时缺失值,Pandas 自动提升列类型为浮点,通过astype(int)可以轻松恢复。

内容的提问来源于stack exchange,提问作者edd313

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 05:05:18