You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何通过单步操作复制或新建Pandas DataFrame并替换指定列值?

嘿,这两个Pandas的小操作我熟,给你整理了清晰的解决方案,结合你给的学生数据例子来演示:

首先先明确咱们的初始DataFrame长啥样(就是你提供的代码跑出来的结果):

student_id    extra_junk   gpa
0    abc123          whoa  3.1
1    def321           hey  junk
2    qwe098  don't touch me  NaN
3    rty135          junk  2.75

1. 单步操作复制DataFrame并替换指定列数值

如果想保留原DataFrame不动,同时生成一个替换了指定列数值的副本,咱们可以用复制+链式替换的单步操作搞定:

import pandas as pd
import numpy as np

# 你的原始数据准备(保留不变)
student_ids = ['abc123', 'def321', 'qwe098', 'rty135']
extra_junk = ['whoa', 'hey', 'don\'t touch me', 'junk']
gpas = ['3.1', 'junk', 'NaN', '2.75']
aa = np.array([student_ids, extra_junk, gpas]).transpose()
df = pd.DataFrame(data=aa, columns=['student_id', 'extra_junk', 'gpa'])

# 核心单步操作:复制+替换指定列
df_copy_replaced = df.copy().replace(
    {'gpa': {'junk': np.nan, 'NaN': np.nan}}
).astype({'gpa': float})

print(df_copy_replaced)

运行后得到的结果:

student_id    extra_junk   gpa
0    abc123          whoa  3.10
1    def321           hey   NaN
2    qwe098  don't touch me   NaN
3    rty135          junk  2.75

这里的关键点:

  • df.copy()确保原DataFrame不受任何修改;
  • replace()的嵌套字典参数精准指定只对gpa列进行替换;
  • 最后用astype()把gpa列从字符串转成数值型,方便后续计算。

2. 单条语句创建新DataFrame并替换指定列值

如果不想先创建原始df再处理,咱们可以把「DataFrame创建」和「列值替换」合并成一条语句完成,推荐两种实用方法:

方法一:链式调用replace+astype

import pandas as pd
import numpy as np

student_ids = ['abc123', 'def321', 'qwe098', 'rty135']
extra_junk = ['whoa', 'hey', 'don\'t touch me', 'junk']
gpas = ['3.1', 'junk', 'NaN', '2.75']

# 单条语句完成创建+替换
new_df = pd.DataFrame(
    data=np.array([student_ids, extra_junk, gpas]).transpose(),
    columns=['student_id', 'extra_junk', 'gpa']
).replace({'gpa': {'junk': np.nan, 'NaN': np.nan}}).astype({'gpa': float})

print(new_df)

方法二:用pd.to_numeric自动处理脏数据

这个方法更简洁,利用pd.to_numeric()的errors='coerce'参数,直接把所有无法转成数值的字符串(比如'junk'、'NaN')自动转为缺失值:

import pandas as pd
import numpy as np

student_ids = ['abc123', 'def321', 'qwe098', 'rty135']
extra_junk = ['whoa', 'hey', 'don\'t touch me', 'junk']
gpas = ['3.1', 'junk', 'NaN', '2.75']

# 更简洁的单条语句
new_df = pd.DataFrame(
    data=np.array([student_ids, extra_junk, gpas]).transpose(),
    columns=['student_id', 'extra_junk', 'gpa']
).assign(gpa=lambda x: pd.to_numeric(x['gpa'], errors='coerce'))

print(new_df)

两种方法得到的结果完全一致,按需选择就行~

内容的提问来源于stack exchange,提问作者Lorem Ipsum

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:28:25