为何SettingWithCopyWarning警告会出现看似不一致的触发情况?
环境准备
我们首先导入pandas和numpy,然后创建一个5000行×10列的pd.DataFrame,命名为df
import pandas as pd import numpy as np # Generate non-trivial df data = np.random.rand(5000, 10) column_names = [f'col_{i}' for i in range(10)] df = pd.DataFrame(data, columns=column_names)
触发SettingWithCopyWarning警告
我们选取df的2列,并为其添加一个值为10的新列
foo = df print(f"foo has type {type(foo)}\n") bar = foo[["col_0", "col_1"]] bar["new"] = 10 print(f"After modification bar has {type(bar)}\n\nbar is:\n{bar}")
执行后得到如下输出:
foo has type <class 'pandas.core.frame.DataFrame'> After modification bar has <class 'pandas.core.frame.DataFrame'> bar is: col_0 col_1 new 0 0.837911 0.715060 10 ... ... ... ... 4999 0.684007 0.176949 10 [5000 rows x 3 columns] C:\Users\vents\AppData\Local\Temp\ipykernel_5824\1295636698.py:5: SettingWithCopyWarning: A value is trying to be set on a copy of a slice from a DataFrame. Try using .loc[row_indexer,col_indexer] = value instead See the caveats in the documentation: https://pandas.pydata.org/pandas-docs/stable/user_guide/indexing.html#returning-a-view-versus-a-copy bar["new"] = 10
注意底部出现的警告信息。
对另一DataFrame执行相同操作却未触发警告
现在我们执行相同操作,但将foo设置为一个小型表格
foo = pd.DataFrame([["x",20.2,30], ["y",80,70]], columns=["col_0", "col_1", "col_2"]) print(f"foo has type {type(foo)}\n") bar = foo[["col_0", "col_1"]] bar["new"] = 10 print(f"After modification bar has {type(bar)}\n\nbar is:\n{bar}")
执行后得到如下输出:
foo has type <class 'pandas.core.frame.DataFrame'> After modification bar has <class 'pandas.core.frame.DataFrame'> bar is: col_0 col_1 new 0 x 20.2 10 1 y 80.0 10
此次未出现警告。
我的猜想与疑问
据我理解,SettingWithCopyWarning意味着无法确定当前表格是copy(副本)还是view(视图)。或许pandas存在优化逻辑:当表格较小时直接返回副本,避免用户困惑;而表格较大时为节省内存,要求用户明确操作?
这是我的初步猜想。
我不愿直接采用通用建议(如使用foo.loc[:, ["col_0", "col_1"]]或foo[["col_0", "col_1"]].copy()来消除警告),因为我尚未完全理解问题本质...
比如我的代码中还有哪些不良写法,只是侥幸未触发警告?
根据网上资料,多数SettingWithCopyWarning会在选取与原表格类型不同的数据子集并对其操作时触发(例如foo[0]["col_0"],先选取序列foo[0]再从中获取列col_0)。
内容的提问来源于stack exchange,提问作者Ventsi Radev
相关产品推荐
相关产品推荐

