如何基于DataFrame其他列生成条件字符串列?
Pandas生成Recipe列的apply方法错误修复
问题场景
现有笛卡尔积结构的DataFrame:
import pandas as pd soup = pd.DataFrame(data={'Beets': [ 1, 2, 0, 1, 2, 0, 1, 2, 0, 1, 2, 0, 1, 2, 0, 1, 2, 0, 1, 2, 0, 1, 2, 0, 1, 2], 'Carrots': [ 0, 0, 1, 1, 1, 2, 2, 2, 0, 0, 0, 1, 1, 1, 2, 2, 2, 0, 0, 0, 1, 1, 1, 2, 2, 2], 'Potatoes': [ 0, 0, 0, 0, 0, 0, 0, 0, 1, 1, 1, 1, 1, 1, 1, 1, 1, 2, 2, 2, 2, 2, 2, 2, 2, 2]})
需要新增Recipe列,规则是:仅保留非零食材,格式为数字+字母代号(Beets→B,Carrots→C,Potatoes→P),多个食材用, 分隔,例如1份Beets、0份Carrots、2份Potatoes对应1B, 2P。
错误原因分析
第一个错误:ValueError: The truth value of a Series is ambiguous.
原代码存在两处核心问题:
- apply用法错误:直接将整个Series传入函数,相当于把Series的计算结果传给apply,而非让apply逐行处理数据。
- 布尔Series判断歧义:函数内
if b>0的判断中,b是整个Series,返回的是布尔Series,if无法直接识别整个Series的真值(不知道取any还是all),导致歧义错误。 - 方法调用错误:错误使用
append[],列表的append是方法,必须用圆括号append()。
第二个错误:KeyError: 'Beets'
修改后的代码未指定axis=1,apply默认按列(axis=0)处理数据,此时传入函数的s是单列的Series,而非行数据,自然无法通过s['Beets']访问列名,触发KeyError。
正确解决方案
修正代码如下,核心是指定axis=1让apply逐行处理,同时修正append的调用方式:
def recipe_namer(row): name = [] if row['Beets'] > 0: name.append(f"{row['Beets']}B, ") if row['Carrots'] > 0: name.append(f"{row['Carrots']}C, ") if row['Potatoes'] > 0: name.append(f"{row['Potatoes']}P, ") # 暂不处理末尾逗号,直接拼接列表元素 return ''.join(name) soup['Recipe'] = soup.apply(recipe_namer, axis=1)
补充优化(可选处理末尾逗号)
如果需要去掉末尾多余的, ,可以在return时做处理:
return ''.join(name).rstrip(', ')
内容的提问来源于stack exchange,提问作者Alexander Zorin
相关产品推荐
相关产品推荐

