You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于国家名数组,从Pandas DataFrame字符串列提取匹配国家生成新列?

嘿,这个需求在数据处理里挺常见的,我给你几个实用的方案,根据你的场景选就行:

方案一:用正则提取(简单高效)

如果你的国家列表里没有名称互相包含的情况(比如不会同时出现China和Chinese),直接用Pandas的str.extract就可以快速搞定:

import pandas as pd
import numpy as np

# 你的国家列表和初始DataFrame
country_list = ['Angola', 'Belgium']
df = pd.DataFrame(
    np.array([['A product for Angola', None], ['A product for Belgium', None], ['A product for France', None]]),
    columns=['Product', 'Country']
)

# 把国家列表拼成正则匹配模式,用|分隔表示“匹配任意一个”
match_pattern = '|'.join(country_list)
# 提取匹配的内容,找不到就返回NaN
df['Country'] = df['Product'].str.extract(f'({match_pattern})', expand=False)

print(df)

运行后你会得到预期的结果:

Product  Country
0  A product for Angola   Angola
1  A product for Belgium  Belgium
2   A product for France      NaN
方案二:自定义函数(灵活可控)

要是你的国家名称有重叠(比如列表里同时出现USA和United States),或者你想控制匹配的优先级,那就用apply搭配自定义函数更合适:

def get_country(text, countries):
    # 遍历国家列表,找到第一个出现在文本里的国家名
    for country in countries:
        if country in text:
            return country
    # 没找到就返回空值
    return np.nan

# 对Product列的每一行应用函数
df['Country'] = df['Product'].apply(lambda x: get_country(x, country_list))

比如你想优先匹配更长的名称,只需要把country_list里的长名称放在前面就行,非常灵活。

额外小提示
  • 如果你的Product列有缺失值,记得先填充空字符串,避免报错:
    df['Product'] = df['Product'].fillna('')
    
  • 要是需要不区分大小写的匹配(比如文本里是"angola",列表里是"Angola"),可以这样改:
    • 正则方案:加个忽略大小写的标志
      import re
      df['Country'] = df['Product'].str.extract(f'({match_pattern})', expand=False, flags=re.IGNORECASE)
      
    • 自定义函数方案:把两者都转成小写再判断
      if country.lower() in text.lower():
      

内容的提问来源于stack exchange,提问作者BrunoGG

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 06:22:47