如何在DataFrame列中搜索字符串并创建条件判断新列
在Pandas DataFrame中搜索字符串并创建条件标记列
没问题,我来帮你搞定这两个需求!咱们从你的示例数据出发,一步步实现。
首先,先把你的示例DataFrame整理得更清晰,给列加上名字方便后续操作:
import pandas as pd # 创建示例DataFrame并指定列名 df = pd.DataFrame([['USDCAD exotic option', -100], ['USDSGD vanilla option', -20]], columns=['Product', 'Value'])
需求一:在列中搜索指定字符串
要在DataFrame的列中筛选包含特定字符串的行,Pandas的str.contains()方法是最直接的选择。它支持忽略大小写匹配,让你的搜索更鲁棒。
比如,筛选包含"exotic"的行:
# 忽略大小写搜索包含"exotic"的行 exotic_rows = df[df['Product'].str.contains('exotic', case=False)] print(exotic_rows)
筛选包含"vanilla"的行同理:
vanilla_rows = df[df['Product'].str.contains('vanilla', case=False)] print(vanilla_rows)
需求二:创建新列标记Vanilla/Exotic
我们可以用NumPy的np.where()函数实现矢量化的条件判断,这种方法效率很高,适合处理大数据集。如果产品名称包含"vanilla"就标记为Vanilla,包含"exotic"就标记为Exotic,其余情况可以标记为Unknown(可选)。
import numpy as np # 创建新列"Type",根据Product列的内容标记类型 df['Type'] = np.where( df['Product'].str.contains('vanilla', case=False), 'Vanilla', np.where( df['Product'].str.contains('exotic', case=False), 'Exotic', 'Unknown' # 可选:处理既不是vanilla也不是exotic的情况 ) ) print(df)
运行后你会得到这样的结果:
Product Value Type 0 USDCAD exotic option -100 Exotic 1 USDSGD vanilla option -20 Vanilla
如果你需要更灵活的逻辑(比如复杂的字符串匹配规则),也可以用apply()方法自定义函数:
def get_option_type(product): product_lower = product.lower() if 'vanilla' in product_lower: return 'Vanilla' elif 'exotic' in product_lower: return 'Exotic' else: return 'Unknown' df['Type'] = df['Product'].apply(get_option_type)
这种方法可读性强,但效率不如矢量化的np.where(),适合小数据集或者逻辑复杂的场景。
内容的提问来源于stack exchange,提问作者Ben H
相关产品推荐
相关产品推荐

