从DataFrame提取指定词生成新列并删除无关行的实现求助
解决方案
以下是满足需求的可行代码实现:
import pandas as pd data = { "Company": [["ConsenSys"], ["Cognizant"], ["IBM"], ["IBM"], ["Reddit, Inc"], ["Reddit, Inc"], ["IBM"]], "skills": [ ['services', 'scientist technical expertise', 'databases'], ['datacomputing tools experience', 'deep learning models', 'cloud services'], ['quantitative analytical projects', 'financial services', 'field experience'], ['filesystems server architectures', 'systems', 'statistical analysis', 'data analytics', 'workflows', 'aws cloud services'], ['aws services'], ['data mining statistics', 'statistical analysis', 'aws cloud', 'services', 'data discovery', 'visualization'], ['communication skills experience', 'services', 'manufacturing environment', 'sox compliance'] ] } dff = pd.DataFrame(data) # 定义需要提取的目标关键词 target_words = ['services', 'statistical analysis'] # 生成new_col:提取skills中包含目标词的关键词 dff['new_col'] = dff['skills'].apply( lambda skills_list: [word for word in target_words if any(word in skill for skill in skills_list)] ) # 删除不包含任意目标词的行,重置索引 dff = dff[dff['new_col'].apply(len) > 0].reset_index(drop=True) print(dff)
代码说明
- 生成new_col:通过
apply遍历每一行的skills列表,检查每个目标关键词是否出现在列表的任意元素中(支持子串匹配,比如aws cloud services会匹配到services),将符合条件的关键词收集到new_col中。 - 过滤无效行:筛选出
new_col不为空的行(即至少包含一个目标词),并重置索引。
如果需要精确匹配(仅提取与目标词完全一致的元素),只需将判断条件修改为:
dff['new_col'] = dff['skills'].apply(lambda skills_list: [word for word in target_words if word in skills_list])
内容的提问来源于stack exchange,提问作者hmmmx2
相关产品推荐
相关产品推荐

