DataFrame列匹配列表元素生成标签列异常问题求助
DataFrame列匹配列表元素生成标签列异常问题求助
问题描述
我有一个DataFrame和两个列表,想实现:如果Description列包含list1里的元素,就新增一列Label标记为weather;如果包含list2里的元素,标记为equipment;都不包含则标记为Other。
相关数据如下:
list1 = ['wind','air'] list2 = ['crane','machine'] # 原始DataFrame df = pd.DataFrame({ 'Description': [ 'There was a heavy wind due to cyclone.', 'Pollution hamper the air quality.', 'The machine failure was due to short circuit.', 'The game was called off due to wind.', 'Players played the game very well.', 'the crane operator took the crane to wrong side' ] })
期望输出:
| Description | Label |
|---|---|
| There was a heavy wind due to cyclone. | weather |
| Pollution hamper the air quality. | weather |
| The machine failure was due to short circuit. | equipment |
| The game was called off due to wind. | weather |
| Players played the game very well. | Other |
| the crane operator took the crane to wrong side | equipment |
我尝试了下面的代码,但最终所有行的Label都变成了Other:
df['Label'] = np.where(df['Description'].str.contains('|'.join(list1)),'weather','Other') df['Label'] = np.where(df['Description'].str.contains('|'.join(list2)),'equipment','Other')
问题原因分析
你遇到的问题是因为两次赋值覆盖了之前的结果:
- 第一次
np.where已经给符合list1的行打上了weather,但第二次np.where会重新判断所有行:只有符合list2的行被设为equipment,其他(包括之前的weather行)都被重置为Other,所以最后只剩equipment和Other,原本的weather全被覆盖了。
解决方案
我们需要按优先级依次判断,或者用嵌套的np.where,或者用np.select来处理多条件判断,这样不会覆盖之前的结果。
方案1:嵌套np.where
先判断list1的条件,再在剩余的行里判断list2的条件,最后设为Other:
import numpy as np import pandas as pd list1 = ['wind','air'] list2 = ['crane','machine'] df = pd.DataFrame({ 'Description': [ 'There was a heavy wind due to cyclone.', 'Pollution hamper the air quality.', 'The machine failure was due to short circuit.', 'The game was called off due to wind.', 'Players played the game very well.', 'the crane operator took the crane to wrong side' ] }) # 嵌套np.where实现多条件判断,case=False忽略大小写匹配 df['Label'] = np.where( df['Description'].str.contains('|'.join(list1), case=False), 'weather', np.where( df['Description'].str.contains('|'.join(list2), case=False), 'equipment', 'Other' ) )
方案2:使用np.select(更清晰的多条件写法)
当条件较多时,np.select可读性更好,它接受条件列表和对应的值列表,最后指定默认值:
conditions = [ df['Description'].str.contains('|'.join(list1), case=False), df['Description'].str.contains('|'.join(list2), case=False) ] choices = ['weather', 'equipment'] df['Label'] = np.select(conditions, choices, default='Other')
补充说明
- 添加
case=False参数可以忽略大小写匹配,比如你的最后一行是小写的crane,也能被正确匹配到; - 如果存在同时符合
list1和list2的行(虽然你的示例里没有),np.where和np.select都会优先匹配第一个条件,所以可以根据需求调整条件的顺序。
运行上面的代码后,就能得到你期望的输出结果啦~
备注:内容来源于stack exchange,提问作者Aditya sharma
相关产品推荐
相关产品推荐

