清洗Pandas DataFrame时遇TypeError:str与int无法比较
Pandas数据清洗:字符串与整数比较的TypeError解决
问题场景
在清洗数据集时需完成以下操作:
- 删除1900年之前的条目
- 删除非电影类型条目
- 删除成人影片条目
执行startYear < 1900判断时触发TypeError,提示字符串与整数无法比较——尽管startYear列内容均为数字,但被识别为字符串类型。
原代码
import pandas as pd df = pd.read_csv('/content/data_mining/merged_dataset.csv') #Cleaning data frame #Delete rows where year is less than 1900: for x in df.index: if df.loc[x, 'startYear'] < 1900: df.drop(x, inplace = True) #Delete rows that are not movies: for x in df.index: if df.loc[x, 'titleType'] != 'movie': df.drop(x, inplace = True) #Delete all adult films: for x in df.index: if df.loc[x, 'isAdult'] == '1': df.drop(x, inplace = True) print(df.to_string())
报错信息
TypeError Traceback (most recent call last) <ipython-input-23-672b76e68b2d> in <cell line: 4>() 3 #Delete rows where year is less than 1900: 4 for x in df.index: ----> 5 if df.loc[x, 'startYear'] < 1900: 6 df.drop(x, inplace = True) 7 TypeError: '<' not supported between instances of 'str' and 'int'
解决方案
1. 转换startYear为数值类型
先将字符串类型的年份转为数值类型,同时处理可能的异常值:
# 转换为整数,无法转换的内容设为NaN df['startYear'] = pd.to_numeric(df['startYear'], errors='coerce')
2. 用矢量化操作高效清洗
Pandas循环遍历索引效率极低,推荐用布尔索引一次性筛选符合条件的行:
import pandas as pd df = pd.read_csv('/content/data_mining/merged_dataset.csv') # 转换startYear为数值类型 df['startYear'] = pd.to_numeric(df['startYear'], errors='coerce') # 同时应用三个筛选条件 df_cleaned = df[ (df['startYear'] >= 1900) & (df['titleType'] == 'movie') & (df['isAdult'] != '1') ].copy() print(df_cleaned.to_string())
优化补充
- 若
isAdult列后续需频繁使用,可转为布尔类型简化操作:df['isAdult'] = df['isAdult'].astype(bool) # 筛选成人影片时可写成 ~df['isAdult'],更直观 - 读取文件时可提前指定列类型(需确保列中无异常值):
df = pd.read_csv( '/content/data_mining/merged_dataset.csv', dtype={'startYear': int, 'isAdult': str} )
内容的提问来源于stack exchange,提问作者Rayed
相关产品推荐
相关产品推荐

