Pandas基于列子串条件生成新列出错,如何修正?
问题解决:Pandas根据子串生成新列全为同一值的问题
你的代码问题出在条件判断的逻辑写法错误,导致所有行都触发第一个if分支,返回"string"。
原代码里的if "a" or "b" or "c" in x: 等价于if True or True or ("c" in x)——因为非空字符串在Python中布尔值为True,所以不管x的内容是什么,这个条件永远成立。
修正后的自定义函数写法
把条件改为每个子串单独判断是否存在于x中:
def function(x): if "a" in x or "b" in x or "c" in x: return "string" elif "d" in x or "e" in x or "f" in x: return "other string" else: return "default string" df['new col'] = df['col'].apply(function) print(df)
更高效的Pandas原生实现(推荐)
避免使用apply(大数据量下效率较低),改用str.contains结合numpy.select实现:
import numpy as np # 定义判断条件和对应返回值 conditions = [ df['col'].str.contains('a|b|c', na=False), df['col'].str.contains('d|e|f', na=False) ] choices = ["string", "other string"] # 生成新列,未匹配到任何条件时返回默认值 df['new col'] = np.select(conditions, choices, default="default string") print(df)
其中str.contains里的a|b|c是正则表达式,代表匹配任意一个目标子串;na=False用于处理空值,避免生成NaN。
内容的提问来源于stack exchange,提问作者MC5555
相关产品推荐
相关产品推荐

