Pandas按其他列设置列值:数据分类报错问题求助
Pandas DataFrame分类处理:拆分 vs 标记及常见错误解决
问题描述
我是Python新手,手里有个从Excel读来的不足1000行的Pandas DataFrame,Store列包含四类数据:门店编号、银行手续费、其他交易、异常数据。纠结是把数据按类型拆成多个DataFrame,还是加分类标记列区分。尝试过程中碰到不少错误,比如str和int比较的TypeError、DataFrame真值模糊的ValueError,附上示例数据和尝试的代码,求正确处理方法。
示例数据及初始化代码
import pandas as pd import numpy as np df = pd.DataFrame.from_dict( { 'Store': ['Bank Fees', 'Bt12600', 'Bt12300', 'Something Else', 'AZ1001', 'TX2002','GA5009'], 'Bank Acct': ['B12343', 'B12344', '', 'B12345','', '', 'B1238'], 'Amount': [1000.00, 2000.00, 1500.00, 2500.00, 55.00, 3000.00, 3500.00], } ) df['Store Length'] = df['Store'].apply(len) df['Store Length'] = df['Store Length'].apply(str) # 为筛选门店转成字符串 df = df.replace('', np.nan) # 空值替换为NaN df['Included'] = "No" # 默认标记为不包含 df['Category'] = 'Exception' # 默认分类为异常 print(df)
尝试代码及报错
- 类型不匹配错误
cond1 = (df.loc[(df['Store Length'].isin(['6','7']))]) cond2 = (df.loc[(df['Bank Acct'].isna())]) # 另一种尝试: df.loc[(df['Store Length'] <= 8) & (df['Bank Acct'].isna())]
报错:TypeError: '<=' not supported between instances of 'str' and 'int'
- 真值模糊错误
if ((cond1) and (cond2)): # 筛选门店:长度<8且Bank Acct为空 df['Category'] = 'STORE' df['Included'] = 'YES' pass
报错:ValueError: The truth value of a DataFrame is ambiguous. Use a.empty, a.bool(), a.item(), a.any() or a.all().
解决方案
一、方案选择:拆分 vs 加标记
数据量不到1000行,两种方案都可行,但加分类标记列更实用:
- 标记列能让后续统计、筛选操作更灵活,不用来回切换多个DataFrame;
- 如果之后需要拆分,也能通过标记快速拆分(比如
df_store = df[df['Category'] == 'STORE']); - 拆分仅适合后续对不同类型数据做完全独立的复杂处理,日常场景用标记列更省心。
二、逐个解决代码错误
1. 修复TypeError:字符串与整数比较
你把Store Length转成了字符串类型,导致和整数8比较报错。没必要转字符串,直接保留整数类型即可:
# 删除转字符串的代码,保留整数长度 df['Store Length'] = df['Store'].apply(len)
之后就能正常用df['Store Length'] <=8做数值比较。
2. 修复ValueError:DataFrame真值模糊
cond1和cond2是筛选后的DataFrame,不是布尔索引,用if判断整个DataFrame的真值会触发歧义报错。正确做法是用布尔索引直接定位行,然后批量赋值,不需要if语句:
# 构建布尔条件 store_condition = (df['Store Length'] <=8) & (df['Bank Acct'].isna()) # 给符合条件的行赋值 df.loc[store_condition, ['Category', 'Included']] = ['STORE', 'YES']
三、完整分类处理代码
结合四类数据的特征,用批量赋值完成全部分类标记:
import pandas as pd import numpy as np df = pd.DataFrame.from_dict( { 'Store': ['Bank Fees', 'Bt12600', 'Bt12300', 'Something Else', 'AZ1001', 'TX2002','GA5009'], 'Bank Acct': ['B12343', 'B12344', '', 'B12345','', '', 'B1238'], 'Amount': [1000.00, 2000.00, 1500.00, 2500.00, 55.00, 3000.00, 3500.00], } ) # 1. 预处理:空值替换+计算Store长度(保留整数) df = df.replace('', np.nan) df['Store Length'] = df['Store'].apply(len) # 2. 初始化标记列 df['Included'] = "No" df['Category'] = 'Exception' # 3. 定义各类数据的布尔条件 bank_condition = df['Store'] == 'Bank Fees' store_condition = (df['Store Length'] <=8) & (df['Bank Acct'].isna()) other_condition = df['Store'] == 'Something Else' # 4. 批量更新分类和标记 df.loc[bank_condition, ['Category', 'Included']] = ['BANK', 'YES'] df.loc[store_condition, ['Category', 'Included']] = ['STORE', 'YES'] df.loc[other_condition, ['Category', 'Included']] = ['OTHER', 'YES'] print(df)
运行后会正确标记每一行的分类,且不会出现之前的错误。
内容的提问来源于stack exchange,提问作者Fred
相关产品推荐
相关产品推荐

