Pandas DataFrame数组列如何精确匹配3个列表元素生成对应新列
Pandas全词匹配提取关键词实现方案
核心需要满足精确全词匹配规则,禁止子串误匹配,通过正则单词边界实现匹配逻辑即可,完整流程如下:
1. 导入依赖、初始化样例数据
import pandas as pd import re import numpy as np # 构造样例数据集 df = pd.DataFrame({ 'id': [1,2,3,5,6,7,8], 'desxription': [ ['this is bad', 'summerfull'], ['city tehran, country iran'], ['uA is a country', 'winternice'], ['this, is, summer'], ['this is winter','uAsal'], ['this is canada' ,'great'], ['this is toronto'] ] }) # 预定义关键词列表 L1 = ['summer', 'winter', 'fall'] L2 = ['iran', 'uA'] L3 = ['tehran', 'canada', 'toronto']
2. 编写通用匹配函数
函数接收单行的描述数组、对应关键词列表,返回第一个匹配到的关键词,无匹配则返回空值:
def match_exact_word(desc_list, keyword_list): # 拼接当前行所有描述文本,空格分隔避免跨字符串连字误判 full_content = ' '.join(desc_list) for kw in keyword_list: # 用单词边界\b包裹关键词,re.escape转义关键词内可能存在的正则特殊字符 match_pattern = re.compile(rf'\b{re.escape(kw)}\b') if match_pattern.search(full_content): return kw return np.nan
3. 批量生成三列结果
对每一列关键词分别调用匹配函数赋值即可:
df['L1'] = df['desxription'].apply(lambda x: match_exact_word(x, L1)) df['L2'] = df['desxription'].apply(lambda x: match_exact_word(x, L2)) df['L3'] = df['desxription'].apply(lambda x: match_exact_word(x, L3))
匹配规则说明
- 正则
\b代表单词边界,只会匹配独立完整的词,既不会把summerfull里的summer误判为命中,也能正确识别tehran,这类带标点的独立词 - 关键词匹配顺序和列表内顺序一致,如果一行存在多个同列表匹配关键词,会返回排在列表前面的结果
- 拼接文本时加空格分隔,避免两个相邻字符串的首尾字符拼接后生成非预期的词造成误判
最终输出结果
执行代码后得到的DataFrame和预期完全一致:
id desxription L1 L2 L3 0 1 [this is bad, summerfull] NaN NaN NaN 1 2 [city tehran, country iran] NaN iran tehran 2 3 [uA is a country, winternice] NaN uA NaN 3 5 [this, is, summer] summer NaN NaN 4 6 [this is winter, uAsal] winter NaN NaN 5 7 [this is canada, great] NaN NaN canada 6 8 [this is toronto] NaN NaN toronto
内容的提问来源于stack exchange,提问作者user15649753
相关产品推荐
相关产品推荐

