Pandas报错:Series的真值判断模糊,求基于行号生成新列的解决方法
问题原因
你遇到的truth value of a Series is ambiguous错误,是因为代码尝试将**单行数据(Series)与多行数据子集(DataFrame)**直接比较,这种比较会返回布尔型DataFrame,而Pandas无法将整个DataFrame的布尔值判断为单一的True/False,导致if语句无法执行。
另外,你的逻辑是想根据行号分类,但代码里却在比较行内容是否等于某子集的内容,完全偏离了目标——你应该检查的是当前行的索引(行号),而非行数据。
修复方案
方案1:利用行索引修改apply函数
修改函数,直接通过行的索引(行号)判断分类:
import pandas as pd shop = pd.DataFrame.from_dict(data) def cate_shop(row_index): if row_index < 4: return 'Service' elif 4 <= row_index < 140: return 'Food & Beverage' elif 140 <= row_index < 173: return 'Fashion' elif 173 <= row_index < 197: return 'Electronics' return 'Other' # 通过row.name获取当前行的索引 shop['category'] = shop.apply(lambda row: cate_shop(row.name), axis=1)
方案2:使用np.select(推荐,效率更高)
对于大型DataFrame,np.select是向量化操作,比逐行apply快得多:
import pandas as pd import numpy as np shop = pd.DataFrame.from_dict(data) # 定义条件列表 conditions = [ shop.index < 4, (shop.index >= 4) & (shop.index < 140), (shop.index >= 140) & (shop.index < 173), (shop.index >= 173) & (shop.index < 197) ] # 对应条件的分类结果 choices = [ 'Service', 'Food & Beverage', 'Fashion', 'Electronics' ] # 生成新列,默认值为'Other' shop['category'] = np.select(conditions, choices, default='Other')
方案3:使用pd.cut(适合连续区间分类)
如果行号是连续的0-based索引,用pd.cut可以更简洁地实现区间划分:
import pandas as pd shop = pd.DataFrame.from_dict(data) # 定义区间边界,左闭右开,include_lowest=True确保第一个区间包含0 bins = [0, 4, 140, 173, 197, shop.index.max() + 1] labels = ['Service', 'Food & Beverage', 'Fashion', 'Electronics', 'Other'] shop['category'] = pd.cut(shop.index, bins=bins, labels=labels, include_lowest=True)
内容的提问来源于stack exchange,提问作者Bee
相关产品推荐
相关产品推荐

