检查列表全部元素是否存在于DataFrame列并生成分类标签列
基于文本内容给DataFrame标记分类标签
需求说明
给现有DataFrame新增Label列,标记规则如下:
- 当
list1中的所有元素都出现在Description列的某一行文本时,标记为weather - 当
list2中的所有元素都出现在该行文本时,标记为equipment - 其余情况标记为
Other
相关数据
list1 = ['wind','air'] list2 = ['crane','failure'] # 原始DataFrame import pandas as pd df = pd.DataFrame({ 'Description': [ 'There was a heavy wind and air dust due to cyclone.', 'Pollution hamper the air quality.', 'The machine failure was due to short circuit.', 'The game was called off due to wind.', 'Players played the game very well.', 'the crane operator took the crane to wrong side.This leads to failure operation.' ] })
实现代码
def get_label(text): # 统一转为小写避免大小写匹配问题 text_lower = text.lower() # 检查list1所有元素是否都存在 if all(item in text_lower for item in list1): return 'weather' # 检查list2所有元素是否都存在 if all(item in text_lower for item in list2): return 'equipment' return 'Other' # 新增Label列 df['Label'] = df['Description'].apply(get_label)
预期输出
Label Description There was a heavy wind and air dust due to cyclone. weather Pollution hamper the air quality. Other The machine failure was due to short circuit. Other The game was called off due to wind. Other Players played the game very well. Other the crane operator took the crane to wrong side.This... equipment
内容的提问来源于stack exchange,提问作者Aditya sharma
相关产品推荐
相关产品推荐

