You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

DataFrame列类型为object却报'int'无split属性错误,求原因及解决方法

问题原因

虽然列的整体数据类型标记为object,但object类型的列允许混合存储不同类型的数据——你的这一列里同时存在字符串和整数类型的值。比如部分行是"300 Pages"这类字符串,而另一部分行已经是纯整数(如300)。当lambda函数遍历到整数类型的行时,调用split()方法自然会触发AttributeError,因为整数没有这个属性。

解决方案

这里提供几种实用的处理方式:

方法1:统一转字符串后处理

先把所有值强制转为字符串,再执行拆分操作,最后统一转为整数:

df['Number of Pages'] = df['Number of Pages'].apply(lambda x: str(x).split(' ')[0]).astype(int)

方法2:使用pandas字符串方法(更简洁)

pandas的str系列方法会自动跳过非字符串类型的值(或返回NaN),可以结合combine_first保留原本的整数值:

# 拆分字符串并取第一部分,对非字符串值返回NaN
split_col = df['Number of Pages'].str.split(' ', expand=True)[0]
# 用原列的整数值填充NaN,最后转int
df['Number of Pages'] = split_col.combine_first(df['Number of Pages'].astype(str)).astype(int)

方法3:针对性处理不同类型的行

先筛选出非字符串的行确认情况,再分别处理:

# 查看所有整数类型的行
int_mask = df['Number of Pages'].apply(lambda x: isinstance(x, int))
print(df[int_mask])

# 只处理字符串类型的行
df.loc[~int_mask, 'Number of Pages'] = df.loc[~int_mask, 'Number of Pages'].apply(lambda x: x.split(' ')[0])
# 统一转为int类型
df['Number of Pages'] = df['Number of Pages'].astype(int)

内容的提问来源于stack exchange,提问作者Armonia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:05:26