pandas中view vs copy问题及SettingWithCopy警告处理方法
pandas解码DataFrame字节类数据触发SettingWithCopy警告的处理
问题描述
我尝试通过以下代码对DataFrame中的字节类数据进行utf-8解码:
df = pd.read_sql_table('mytable',con) for column in df.columns : for i in range(len(df[column])): if type(df[column][i]) == bytearray or type(df[column][i]) == bytes: df[column][i] = str(df[column][i], 'utf-8')
但无论怎么调整写法,运行时始终触发SettingWithCopyWarning,需要找到正确的处理方式消除警告、实现解码需求。
最终可行方案
通过判断列类型+整列apply批量处理的方式实现需求,完全避免链式索引触发的警告,代码如下:
if df[column].dtype == 'object': df[column] = df[column].apply(lambda x: x.decode('utf-8') if isinstance(x, bytes) else x)
问题原因说明
- 原始写法触发警告的核心问题是使用了
df[column][i]的链式索引赋值逻辑,pandas无法确认当前操作的是原DataFrame的视图还是独立副本,因此抛出警告提示赋值可能存在不生效的风险。 - 整列
apply的写法直接对列对象做整体替换,不存在链式索引的问题,同时批量处理的运行效率远高于逐行遍历循环的写法。
内容的提问来源于stack exchange,提问作者Pefori
相关产品推荐
相关产品推荐

