如何用for循环替代condition list简化np.select代码?解决参数错误
问题解决:循环实现字段赋值及轻量化方案
报错翻译
first argument must be string or compiled pattern → 第一个参数必须是字符串或编译后的正则表达式
错误原因
- 循环中使用的
i是range()生成的整数,而str.contains()要求传入字符串或正则表达式,类型不匹配导致报错。 - 每次循环直接覆盖
conditionlist,最终仅保留最后一次循环的条件,完全偏离原逻辑。 - 尝试代码中
choicelist里的biurbon是拼写错误,应为bourbon。
正确循环实现方式
先建立关键词(带下划线)与产品类型的映射关系,再通过循环生成条件列表和选择列表,既避免重复代码,又保证逻辑正确:
import numpy as np import pandas as pd # 建立关键词-产品类型映射(对应原代码的条件与选择关系) type_mapping = { 'vin_': 'champagne', 'cremant_': 'vin', 'biere_': 'cremant', 'cidre_': 'biere', 'pastis_': 'cidre', 'whisky_': 'pastis', 'bourbon_': 'whisky', 'brandy_': 'bourbon', 'tequila_': 'brandy', 'vodka_': 'tequila', 'gin_': 'vodka', 'rhum_': 'gin' } # 循环生成条件列表和选择列表 conditionlist = [] choicelist = [] for keyword, product_type in type_mapping.items(): conditionlist.append(df_only_total_sales['post_name'].str.contains(keyword)) choicelist.append(product_type) # 赋值产品类型 df_only_total_sales['Product_type'] = np.select(conditionlist, choicelist, default="?") df_only_total_sales.sample(10)
轻量化优化方案
使用正则提取+字典映射的方式,替代np.select,代码更简洁且执行效率更高:
# 生成匹配所有关键词的正则表达式 pattern = '|'.join(type_mapping.keys()) # 提取post_name中匹配的关键词 df_only_total_sales['matched_keyword'] = df_only_total_sales['post_name'].str.extract(f'({pattern})', expand=False) # 通过映射直接得到产品类型,默认值为"?" df_only_total_sales['Product_type'] = df_only_total_sales['matched_keyword'].map(type_mapping).fillna("?") # 可选:删除中间列 df_only_total_sales.drop('matched_keyword', axis=1, inplace=True) df_only_total_sales.sample(10)
内容的提问来源于stack exchange,提问作者Diastat
相关产品推荐
相关产品推荐

