You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Pandas pipe增列时原地修改DataFrame,删改行列不生效?如何解决?

问题

在使用Pandas的pipe方法时发现:调用添加列的函数会原地修改原DataFrame,但调用删除行/列、添加行的函数时,原DataFrame不会被修改。请问这是预期行为吗?如何无需重新赋值变量就能确保DataFrame始终被原地修改?

测试代码及输出如下:

# Dummy data
tdf = pd.DataFrame(dict(a=[1, 2, 3, 4, 5], b=[33, 22, 66, 33, 77]))

# Functions to pipe
def addcol(dataf):
    dataf["c"] = 1000
    return dataf

def remcol(dataf):
    dataf = dataf.drop(columns='c')
    return dataf

def addrow(dataf):
    dataf = pd.concat([dataf, dataf])
    return dataf

def remrow(dataf):
    dataf = dataf.loc[dataf.a < 4]
    return dataf

# Utility function to print result
def printer(pdf, funcname):
    shape1 = pdf.pipe(eval(funcname)).shape
    shape2 = pdf.shape
    if (shape1 == shape2):
        print(f"{funcname}: DataFrame updated in place: shape1 = {shape1}, shape2 = {shape2}")
    else:
        print(f"{funcname}: DataFrame NOT updated in place: shape1 = {shape1}, shape2 = {shape2}")

for fn in ["addcol", "remcol", "addrow", "remrow"]:
    printer(tdf, fn)

运行输出:

addcol: DataFrame updated in place: shape1 = (5, 3), shape2 = (5, 3)
remcol: DataFrame NOT updated in place: shape1 = (5, 2), shape2 = (5, 3)
addrow: DataFrame NOT updated in place: shape1 = (10, 3), shape2 = (5, 3)
remrow: DataFrame NOT updated in place: shape1 = (3, 3), shape2 = (5, 3)

使用Pandas版本:2.0.0

解答

这是预期行为吗?

是预期行为,核心差异来自Pandas操作的两种类型:

  • 原地修改操作:像dataf["c"] = 1000这类直接操作原DataFrame内存空间的行为,不会创建新对象,所以原DataFrame会被同步修改。
  • 返回新对象操作:drop()、concat()、loc[]这类方法默认会生成新的DataFrame对象,你代码里将函数内的dataf变量重新赋值为新对象,只是改变了局部变量的指向,完全不影响原DataFrame的内存空间,所以原对象不会被修改。

如何确保无需重新赋值就能原地修改?

要让所有操作都实现原地修改,需要调整函数逻辑,避免创建新对象,直接操作原DataFrame:

  1. 删除行/列:使用带inplace=True参数的方法(Pandas 2.0+中这类方法返回None,所以函数内不需要重新赋值)
  2. 添加行:通过df.loc直接写入原DataFrame的内存空间,避免使用concat生成新对象

修改后的示例代码:

import pandas as pd

# Dummy data
tdf = pd.DataFrame(dict(a=[1, 2, 3, 4, 5], b=[33, 22, 66, 33, 77]))

# 调整为原地修改的pipe函数
def addcol(dataf):
    dataf["c"] = 1000
    return dataf

def remcol(dataf):
    dataf.drop(columns='c', inplace=True)
    return dataf

def addrow(dataf):
    # 直接用loc追加行,原地修改
    new_rows = dataf.copy()
    dataf.loc[dataf.index.max() + 1 : dataf.index.max() + len(new_rows)] = new_rows.values
    return dataf

def remrow(dataf):
    # 获取保留行的索引,原地删除其余行
    keep_idx = dataf[dataf.a < 4].index
    dataf.drop(dataf.index.difference(keep_idx), inplace=True)
    return dataf

# Utility function to print result
def printer(pdf, funcname):
    shape1 = pdf.pipe(eval(funcname)).shape
    shape2 = pdf.shape
    if (shape1 == shape2):
        print(f"{funcname}: DataFrame updated in place: shape1 = {shape1}, shape2 = {shape2}")
    else:
        print(f"{funcname}: DataFrame NOT updated in place: shape1 = {shape1}, shape2 = {shape2}")

for fn in ["addcol", "remcol", "addrow", "remrow"]:
    printer(tdf, fn)

修改后运行输出:

addcol: DataFrame updated in place: shape1 = (5, 3), shape2 = (5, 3)
remcol: DataFrame updated in place: shape1 = (5, 2), shape2 = (5, 2)
addrow: DataFrame updated in place: shape1 = (10, 2), shape2 = (10, 2)
remrow: DataFrame updated in place: shape1 = (6, 2), shape2 = (6, 2)

注意事项

  • inplace=True操作不可逆,且在链式调用中需注意:因为这类方法返回None,所以函数最后要手动返回原DataFrame才能继续pipe链式调用。
  • 追加行的方式:避免用concat生成新对象,直接通过loc写入原DataFrame是更纯粹的原地修改方式。

内容的提问来源于stack exchange,提问作者Saaru Lindestøkke

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.24 19:05:32