基于非数值标识符拆分DataFrame列内容报错及处理需求
问题解决:AttributeError: 'int' object has no attribute 'isnumeric'
错误原因
你的index1列混合了整数和字符串类型:原行索引是整数,手动添加的小计行索引是字符串。当循环到整数类型的元素时,调用.isnumeric()方法会报错——因为整数对象本身没有这个字符串方法。
修复方案
1. 统一索引列的类型
先把index1列全部转为字符串类型,确保所有元素都能调用字符串方法:
df['index1'] = df.index.astype(str)
2. 用矢量化操作替代for循环(更高效)
Pandas不推荐用for循环遍历行,改用矢量化的字符串操作处理,同时注意分类类型的兼容性:
# 扩展Source列的分类,避免赋值时出现NaN df['Source'] = df['Source'].cat.add_categories(['In Progress Subtotal']) # 筛选出非数值的索引行 mask = ~df['index1'].str.isnumeric() # 拆分字符串并赋值给对应列(strip()去除拆分后的空格) split_cols = df.loc[mask, 'index1'].str.split('-', expand=True) df.loc[mask, 'Status'] = split_cols[0].str.strip() df.loc[mask, 'Source'] = split_cols[1].str.strip() # 可选:不需要index1列可删除 # df.drop('index1', axis=1, inplace=True)
完整修复后的代码
import pandas as pd df = pd.DataFrame( {'Status':['Active','Active','Inactive','Active'], 'Source':['Domestic','International','International','Restricted'], 'Activity':['In Progress','Post','Post','In Progress'], 'FY20':[2,55,52,99],'FY21':[90,20,11,43],'FY22':[10,52,57,9]}) df['Status'] = pd.Categorical(df['Status'],categories= ['Inactive','Active']) df = df.sort_values('Status') df['Source'] = pd.Categorical(df['Source'],categories= ['International','Restricted','Domestic']) df = df.sort_values('Source') # 添加小计行 df.loc['Active - In Progress Subtotal'] = df.loc[(df['Status'].isin(['Active'])) & (df['Activity'].isin(['In Progress']))].sum(axis=0) # 转换索引列为字符串 df['index1'] = df.index.astype(str) # 处理拆分逻辑 df['Source'] = df['Source'].cat.add_categories(['In Progress Subtotal']) mask = ~df['index1'].str.isnumeric() split_cols = df.loc[mask, 'index1'].str.split('-', expand=True) df.loc[mask, 'Status'] = split_cols[0].str.strip() df.loc[mask, 'Source'] = split_cols[1].str.strip() # 查看结果 print(df)
关键说明
- 分类类型(Categorical)必须提前添加新的类别,否则赋值时会因为类别不存在而变成
NaN - 矢量化操作比for循环效率高得多,尤其适合大型DataFrame
.str.strip()是为了去除拆分后字符串两端的空格(比如拆分'Active - In Progress Subtotal'后,第二部分开头有空格)
内容的提问来源于stack exchange,提问作者Sam
相关产品推荐
相关产品推荐

