如何在Pandas中从指定国家列表匹配值替换DataFrame描述列字符串
最优实现方案
推荐使用str.extract矢量化操作实现,比循环遍历国家列表的性能高很多,代码也更简洁:
import re # 拼接正则匹配规则,自动转义国家名中的特殊字符 country_pattern = '|'.join(re.escape(item) for item in country_list) # 忽略大小写提取描述中匹配的国家名,无匹配时保留原描述 df['Description'] = df['Description'].str.extract( f'({country_pattern})', flags=re.IGNORECASE, expand=False ).fillna(df['Description'])
如果需要统一输出国家名的大小写和你提供的国家列表完全一致,再加一步映射即可:
# 构建大小写不敏感的国家映射字典 country_mapping = {c.lower(): c for c in country_list} df['Description'] = df['Description'].str.extract( f'({country_pattern})', flags=re.IGNORECASE, expand=False ).str.lower().map(country_mapping).fillna(df['Description'])
方案优势
- 完全基于pandas原生矢量化运算,无需循环遍历国家列表,数据量越大性能优势越突出
- 仅需一次全量匹配即可完成所有替换,没有重复计算
- 代码简洁易读,符合Pythonic的编码规范
内容的提问来源于stack exchange,提问作者oettam_oisolliv
相关产品推荐
相关产品推荐

