如何在Pandas中基于列表列创建匹配指定字符串的新列
解决方法
步骤1:构造测试数据
先把示例数据转换成Pandas DataFrame:
import pandas as pd data = { 'id': [123, 234, 543, 542, 857, 123], 'specification': [ ['high', 'Important', 'pilot'], ['HIGH', 'Important', 'Baby'], ['important'], ['week'], ['new', 'IMPORTANT'], ['super_high', 'test'] ] } df = pd.DataFrame(data)
步骤2:定义匹配逻辑
用正则实现大小写不敏感的关键词匹配,同时支持检测包含关键词的字符串(比如super_high里的high)。写个辅助函数遍历列表元素,检查是否符合条件:
import re def check_spec(spec_list): # 匹配规则:包含high或important,不区分大小写 pattern = re.compile(r'(high|important)', re.IGNORECASE) # 只要列表中有一个元素匹配,就返回指定值 for item in spec_list: if pattern.search(item): return 'High/Important' # 无匹配项返回空字符串 return ''
步骤3:生成新列
用apply把辅助函数应用到specification列,生成new_col:
df['new_col'] = df['specification'].apply(check_spec)
最终结果
运行后得到的结果和预期一致:
id specification new_col 0 123 [high, Important, pilot] High/Important 1 234 [HIGH, Important, Baby] High/Important 2 543 [important] High/Important 3 542 [week] 4 857 [new, IMPORTANT] High/Important 5 123 [super_high, test] High/Important
补充说明
- 正则表达式
r'(high|important)'搭配re.IGNORECASE,能匹配任意大小写的目标关键词,还能识别包含这些关键词的衍生字符串(比如super_high、IMPORTANT_LEVEL)。 - 要是需要扩展关键词,直接修改正则内容就行,比如加
critical就改成r'(high|important|critical)'。
内容的提问来源于stack exchange,提问作者JackJack
相关产品推荐
相关产品推荐

