You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Pandas中基于多字符串条件生成分类标记列?

Pandas实现国家列的US区域标记

你已经找到了正确的起点:用str.contains生成布尔序列来判断行中是否包含目标国家,只需要将这个布尔序列转换为你需要的字符串标记即可。以下是几种高效实现方式:

方法一:使用numpy.where(推荐,向量化操作效率高)

np.where可以根据布尔条件直接返回对应的值,是处理这类映射最常用的方法:

import pandas as pd
import numpy as np

# 构造测试DataFrame
df = pd.DataFrame({
    'countries': [
        'United States, Japan',
        'China',
        'Brazil, South Africa',
        'Puerto Rico, Spain',
        'United States, Vietnam',
        'Madagascar'
    ]
})

# 生成标记列
df['region'] = np.where(
    df['countries'].str.contains('United States|Puerto Rico'),
    'Inside US',
    'Outside US'
)

# 查看结果
print(df['region'])

输出结果:

0    Inside US
1    Outside US
2    Outside US
3    Inside US
4    Inside US
5    Outside US
Name: region, dtype: object

方法二:使用Series.map转换布尔值

直接将布尔序列映射为目标字符串,代码更简洁:

df['region'] = df['countries'].str.contains('United States|Puerto Rico').map(
    {True: 'Inside US', False: 'Outside US'}
)

方法三:使用apply逐行判断(适合逻辑复杂场景)

如果后续需要扩展判断逻辑,可以用apply自定义函数:

def get_region(country_text):
    if 'United States' in country_text or 'Puerto Rico' in country_text:
        return 'Inside US'
    return 'Outside US'

df['region'] = df['countries'].apply(get_region)

方法对比

  • 方法一和方法二属于向量化操作,处理大数据量时性能远优于逐行处理的apply,优先推荐。
  • 方法三代码逻辑直观,但数据量较大时效率较低,适合逻辑复杂、数据量小的场景。

内容的提问来源于stack exchange,提问作者Eisen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 14:35:22