Pandas创建条件新列时消除链式赋值警告的方案求助
问题
我写了一段用双条件创建新列的Pandas脚本,功能正常,但触发了ChainedAssignmentError(未来警告,Pandas 3.0行为将变更)和SettingWithCopyWarning。以下是代码和报错信息,求优化消除警告的方法:
import pandas as pd import numpy as np import matplotlib.pyplot as plt df=pd.DataFrame() df['variable 1']= np.arange(0,1.1,0.1) df['variable 2']= 0.2*df['variable 1'] df['variable 3']= 0.4 -0.2*df['variable 1'] # Create new columns slope = [2, 1.5, 1, 0.5] for i in range(len(slope)): df['slope = ' + str(slope[i])]='' for j in range(len(df['variable 1'])): # Calculating Scl_disp_sd with equation 1 curve = 0.5 - slope[i]*df['variable 1'][j] df['slope = ' + str(slope[i])][j]= np.where((curve>df['variable 2'][j]) & (curve<df['variable 3'][j]), curve,np.nan) display(df) plt.plot(df['variable 1'], df['variable 2'], 'o', label='variable 2') plt.plot(df['variable 1'], df['variable 3'], 'o', label='variable 3') plt.plot(df['variable 1'], df.filter(like='slope =', axis=1), marker='.') plt.legend()
报错信息:
/var/folders/m0/_y1fs5x50xx99pjg2yf42y7r0000gp/T/ipykernel_1964/2618301266.py:11: FutureWarning: ChainedAssignmentError: behaviour will change in pandas 3.0! You are setting values through chained assignment. Currently this works in certain cases, but when using Copy-on-Write (which will become the default behaviour in pandas 3.0) this will never work to update the original DataFrame or Series, because the intermediate object on which we are setting values will behave as a copy. A typical example is when you are setting values in a column of a DataFrame, like: df["col"][row_indexer] = value Use `df.loc[row_indexer, "col"] = values` instead, to perform the assignment in a single step and ensure this keeps updating the original `df`. See the caveats in the documentation: https://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html#returning-a-view-versus-a-copy df['slope = ' + str(slope[i])][j]= np.where((curve>df['variable 2'][j]) & (curve<df['variable 3'][j]), /var/folders/m0/_y1fs5x50xx99pjg2yf42y7r0000gp/T/ipykernel_1964/2618301266.py:11: SettingWithCopyWarning: A value is trying to be set on a copy of a slice from a DataFrame See the caveats in the documentation: https://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html#returning-a-view-versus-a-copy df['slope = ' + str(slope[i])][j]= np.where((curve>df['variable 2'][j]) & (curve<df['variable 3'][j]), ...
优化方案
核心问题
警告根源是链式索引赋值:df['col'][j]这种写法先取列再取行,中间可能生成副本,导致Pandas无法确认你要修改原DataFrame还是副本。另外,嵌套循环逐行赋值效率极低,违背Pandas向量化操作的设计理念。
优化步骤
- 替换链式索引为单步定位:用
df.loc[j, 'col']直接定位行和列,明确修改原DataFrame(如果保留循环的话)。 - 用向量化操作替代循环:Pandas支持整列批量运算,完全不需要逐行循环,既能消除警告又大幅提升效率。
修改后的代码
import pandas as pd import numpy as np import matplotlib.pyplot as plt df = pd.DataFrame() df['variable 1'] = np.arange(0, 1.1, 0.1) df['variable 2'] = 0.2 * df['variable 1'] df['variable 3'] = 0.4 - 0.2 * df['variable 1'] slope_list = [2, 1.5, 1, 0.5] for s in slope_list: col_name = f'slope = {s}' # 向量化计算整条曲线 curve = 0.5 - s * df['variable 1'] # 双条件过滤,直接生成整列数据 df[col_name] = np.where((curve > df['variable 2']) & (curve < df['variable 3']), curve, np.nan) display(df) plt.plot(df['variable 1'], df['variable 2'], 'o', label='variable 2') plt.plot(df['variable 1'], df['variable 3'], 'o', label='variable 3') plt.plot(df['variable 1'], df.filter(like='slope =', axis=1), marker='.') plt.legend()
关键改进说明
- 移除内层逐行循环,直接对整列运算,数据量越大效率提升越明显。
- 用
df[col_name] = ...直接赋值整列,彻底避免链式索引问题,消除所有警告。 - 使用f-string格式化列名,代码更简洁易读。
内容的提问来源于stack exchange,提问作者Pablo
相关产品推荐
相关产品推荐

