Python Pandas分类列报错:无法执行ror_操作及类型转换失败问题
问题解决方案
一、修复分类编码的循环错误
你的代码主要有两个逻辑错误:
- 误用了位运算符
&/|,应该用逻辑运算符and/or做单条件判断 - 判断列名时错误地拿列数据和字符串对比,正确方式是直接比较列名
col
修正后的代码
df_cat = df.copy() # 复制原数据,避免修改原始DataFrame # 方式1:先筛选目标列再循环 target_cols = [col for col in df_cat.columns if df_cat[col].dtype == 'object' and col not in ['rating', 'correctYear']] for col in target_cols: df_cat[col] = df_cat[col].astype('category') df_cat[col] = df_cat[col].cat.codes
或者直接在循环内修正条件:
df_cat = df.copy() for col in df_cat.columns: # 正确条件:列类型为object,且列名不是rating和correctYear if df_cat[col].dtype == 'object' and col != 'rating' and col != 'correctYear': df_cat[col] = df_cat[col].astype('category') df_cat[col] = df_cat[col].cat.codes
二、解决correctYear转数值的报错
报错could not convert string to float:说明correctYear列存在非数字内容(比如空值、特殊字符、字母),可以按以下方式处理:
1. 强制转换(无法转换的设为NaN)
import pandas as pd df_cat['correctYear'] = pd.to_numeric(df_cat['correctYear'], errors='coerce') # 可选:用year列的值填充NaN df_cat['correctYear'] = df_cat['correctYear'].fillna(df_cat['year'])
2. 先清理非数字内容再转换
如果列中存在类似"2020年"的格式,先提取数字:
df_cat['correctYear'] = df_cat['correctYear'].str.extract('(\d+)').astype(float)
错误原因说明
原代码中的条件(df_cat[col].dtype=='object' & (df_cat[col]!='rating' | df_cat[col]!='correctYear'))有两个问题:
&/|是数组位运算符号,不适合用来做单个布尔条件的逻辑判断,应该用and/ordf_cat[col]!='rating'是把整列数据和字符串对比,得到布尔数组,而不是判断列名是否为rating,这导致了数组与布尔值的运算冲突,触发报错
内容的提问来源于stack exchange,提问作者Rybin Mathalikunnel
相关产品推荐
相关产品推荐

