Pandas查找列中子串写入另一列时报真值歧义错误如何解决?
问题原因及解决方法
你遇到的The truth value of a DataFrame is ambiguous.报错核心原因是:你直接将布尔筛选后的DataFrame放在if判断语句中,pandas无法将包含多行多值的DataFrame转换为单个布尔值,因此抛出该错误。
除此之外你的原有代码还有以下几处问题:
isin()方法用于判断元素是否在指定的可迭代对象中,你传入单个字符串会被拆解为单个字符匹配,完全不符合匹配特定子串的需求。你的场景中Percentage列格式固定为数字单词开头,用str.startswith()匹配更精准。str.replace后你使用了方括号[]调用,属于语法错误,应该用圆括号()。- 直接对整列赋值的写法会覆盖所有行的结果,无法实现按行匹配条件赋值的需求。
最优实现方案
推荐用numpy.select做多条件匹配,是pandas处理多分支条件赋值的标准写法,执行效率远高于逐行遍历:
import pandas as pd import numpy as np # 读取CSV文件 df = pd.read_csv(r'file.csv', dtype=str) # 定义条件列表和对应的等级列表 conditions = [ df['Percentage'].str.startswith('Ninety'), df['Percentage'].str.startswith('Eighty'), df['Percentage'].str.startswith('Seventy'), df['Percentage'].str.startswith('Sixty') ] grades = ['A', 'B', 'C', 'D'] # 按条件赋值,不满足所有条件的默认赋值为F df['Letter Grade'] = np.select(conditions, grades, default='F') # 保存回CSV df.to_csv(r'file.csv', index=False)
如果不想引入numpy依赖,也可以用df.loc写法实现:
import pandas as pd df = pd.read_csv(r'file.csv', dtype=str) # 先默认全赋值为F df['Letter Grade'] = 'F' df.loc[df['Percentage'].str.startswith('Ninety'), 'Letter Grade'] = 'A' df.loc[df['Percentage'].str.startswith('Eighty'), 'Letter Grade'] = 'B' df.loc[df['Percentage'].str.startswith('Seventy'), 'Letter Grade'] = 'C' df.loc[df['Percentage'].str.startswith('Sixty'), 'Letter Grade'] = 'D' df.to_csv(r'file.csv', index=False)
内容的提问来源于stack exchange,提问作者FinestRyeBread
相关产品推荐
相关产品推荐

