Python创建RAM分类变量报错求助:TypeError问题修复
解决pd.cut中int与str比较的TypeError问题
问题背景
需要为RAM创建分类变量,划分规则如下:
- Basic:RAM [0-4]
- Intermediate:RAM [5-8]
- Advanced:RAM [8-12]
执行的代码:
df['Memory']=pd.cut(df['RAM '], [0,4,8,12], include_lowest=True, labels=['Basic','Intermediate', 'Advaced'])
运行后出现错误:
TypeError Traceback (most recent call last) <ipython-input-58-5c93d7c00ba2> in <cell line: 1>() ----> 1 df['Memory']=pd.cut(df['RAM '], [0,4,8,12], include_lowest=True, labels=['Basic', 'Intermediate', 'Advaced']) 1 frames /usr/local/lib/python3.9/dist-packages/pandas/core/reshape/tile.py in _bins_to_cuts(x, bins, right, labels, precision, include_lowest, dtype, duplicates, ordered) 425 426 side: Literal["left", "right"] = "left" if right else "right" --> 427 ids = ensure_platform_int(bins.searchsorted(x, side=side)) 428 429 if include_lowest: TypeError: '<' not supported between instances of 'int' and 'str'
错误原因
报错核心是df['RAM ']列的数据类型为字符串(str),而pd.cut需要基于数值类型(int/float)进行区间划分,无法直接对字符串和整数做比较运算。另外代码里存在两处小问题:列名多了末尾空格、分类标签里Advanced拼写错误。
修复步骤
1. 转换RAM列为数值类型
先处理列名和数据类型:
# 修正列名(去掉末尾多余空格) df.rename(columns={'RAM ': 'RAM'}, inplace=True) # 情况1:RAM列是纯数字字符串,直接转整数 df['RAM'] = df['RAM'].astype(int) # 情况2:RAM列带单位(如"4GB"、"8 GB"),先提取数字再转换 df['RAM'] = df['RAM'].str.extract('(\d+)').astype(int)
2. 修正分类代码
调整拼写并确保区间逻辑合理:
df['Memory'] = pd.cut(df['RAM'], [0,4,8,12], include_lowest=True, labels=['Basic','Intermediate', 'Advanced'])
3. 验证结果
执行以下命令确认数据类型正确:
print(df['RAM'].dtype) # 输出应为int64或float64
内容的提问来源于stack exchange,提问作者codproe
相关产品推荐
相关产品推荐

