Pandas两个DataFrame条件匹配 按匹配结果生成指定值新列问题
问题原因
你原有代码的错误逻辑有两点:
- 调用
any()会判断整个lead_investors列是否存在任意一行匹配当前投资者名,只要有一个匹配就返回True - 条件成立时直接对整列
investor_continent赋值'asia',最终只要名单里有任意一个投资者在列中出现过,全列都会被覆盖为'asia'
最优实现(推荐)
不需要循环,用pandas向量化操作一次性完成匹配,效率远高于循环遍历:
import re # 1. 拼接正则匹配模式,re.escape用于转义投资者名称中的正则特殊字符 pattern = '|'.join([re.escape(name) for name in names]) # 2. 生成匹配掩码:每行只要包含名单中任意一个投资者名称,就标记为True match_mask = df1c['lead_investors'].str.contains(pattern, na=False) # 3. 按匹配结果赋值 df1c['investor_continent'] = match_mask.map({True: 'asia', False: 'other'})
循环实现(不推荐,仅作逻辑参考)
如果一定要用循环实现,修改逻辑如下:
# 先默认所有行都为other df1c['investor_continent'] = 'other' for x in names: # 仅筛选匹配当前投资者的行,单独修改这些行的取值 match_rows = df1c['lead_investors'].str.contains(x, na=False) df1c.loc[match_rows, 'investor_continent'] = 'asia'
内容的提问来源于stack exchange,提问作者Steven
相关产品推荐
相关产品推荐

