如何在Pandas中基于多字符串条件生成分类标记列?
Pandas实现国家列的US区域标记
你已经找到了正确的起点:用str.contains生成布尔序列来判断行中是否包含目标国家,只需要将这个布尔序列转换为你需要的字符串标记即可。以下是几种高效实现方式:
方法一:使用numpy.where(推荐,向量化操作效率高)
np.where可以根据布尔条件直接返回对应的值,是处理这类映射最常用的方法:
import pandas as pd import numpy as np # 构造测试DataFrame df = pd.DataFrame({ 'countries': [ 'United States, Japan', 'China', 'Brazil, South Africa', 'Puerto Rico, Spain', 'United States, Vietnam', 'Madagascar' ] }) # 生成标记列 df['region'] = np.where( df['countries'].str.contains('United States|Puerto Rico'), 'Inside US', 'Outside US' ) # 查看结果 print(df['region'])
输出结果:
0 Inside US 1 Outside US 2 Outside US 3 Inside US 4 Inside US 5 Outside US Name: region, dtype: object
方法二:使用Series.map转换布尔值
直接将布尔序列映射为目标字符串,代码更简洁:
df['region'] = df['countries'].str.contains('United States|Puerto Rico').map( {True: 'Inside US', False: 'Outside US'} )
方法三:使用apply逐行判断(适合逻辑复杂场景)
如果后续需要扩展判断逻辑,可以用apply自定义函数:
def get_region(country_text): if 'United States' in country_text or 'Puerto Rico' in country_text: return 'Inside US' return 'Outside US' df['region'] = df['countries'].apply(get_region)
方法对比
- 方法一和方法二属于向量化操作,处理大数据量时性能远优于逐行处理的
apply,优先推荐。 - 方法三代码逻辑直观,但数据量较大时效率较低,适合逻辑复杂、数据量小的场景。
内容的提问来源于stack exchange,提问作者Eisen
相关产品推荐
相关产品推荐

