You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何修正Pandas DataFrame按project_ids替换hours时的NaN问题?有更优方案吗?

问题原因分析

你遇到的NaN问题,根源在于索引不匹配:当你用replacements.loc[replacements['project_ids'] == project, 'hours']取值时,返回的是一个带索引的Series对象,而非单个标量值。比如:

  • 当project=2时,replacements中对应的行索引是0,而df中project=2的行索引是1,赋值时因为索引无法对齐,就会填充NaN。
  • 巧合的是project=3在replacements和df中的行索引都是2,索引刚好匹配,所以替换成功了。

修复原有循环的方法

只需要把replacements取出的Series转成标量值即可,比如用.iloc[0]或者.values[0]来获取单个值:

import pandas as pd
df = pd.DataFrame({
    'project_ids': [1, 2, 3, 4, 5],
    'hours': [111, 222, 333, 444, 555],
    'else' :['a', 'b', 'c', 'd', 'e']
})
replacements = pd.DataFrame({
    'project_ids': [2, 5, 3],
    'hours': [666, 999, 1000],
})

for project in replacements['project_ids']:
    # 用.iloc[0]获取标量值
    new_hours = replacements.loc[replacements['project_ids'] == project, 'hours'].iloc[0]
    df.loc[df['project_ids'] == project, 'hours'] = new_hours

print(df)

运行后就能得到正确结果:

project_ids  hours else
0           1    111    a
1           2    666    b
2           3   1000    c
3           4    444    d
4           5    999    e

更优的实现方式(推荐)

Pandas的核心优势是矢量化操作,循环在数据量大时效率极低,推荐以下几种更高效的方法:

方法1:使用pd.Series.map()

先把replacements转成字典,再用map匹配替换,不匹配的保留原数值:

# 构建project_id到hours的映射字典
hour_map = replacements.set_index('project_ids')['hours'].to_dict()
# 替换:如果project_id在字典中就用对应值,否则保留原hours
df['hours'] = df['project_ids'].map(hour_map).fillna(df['hours'])

方法2:使用DataFrame.update()

先把replacements设置索引为project_ids,然后用update直接替换匹配的行:

# 把replacements的索引设为project_ids,只保留hours列
replacements_indexed = replacements.set_index('project_ids')['hours']
# update会根据索引匹配替换,只替换存在匹配的行
df.set_index('project_ids', inplace=True)
df.update(replacements_indexed)
df.reset_index(inplace=True)

方法3:使用DataFrame.merge()

通过合并找到匹配项,再合并回原DataFrame:

# 合并两个DataFrame,只保留需要替换的hours列
merged = df.merge(replacements, on='project_ids', how='left', suffixes=('', '_new'))
# 用新的hours值替换原hours,没有匹配的保留原数值
df['hours'] = merged['hours_new'].fillna(merged['hours'])
# 删掉临时列(可选)
df.drop('hours_new', axis=1, inplace=True)

这几种方法都比循环高效,尤其是当数据量较大时,性能差距会非常明显。

内容的提问来源于stack exchange,提问作者barciewicz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:34:25