You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从DataFrame每行文本中提取指定技能词汇生成新列

提取DataFrame文本列中所有匹配技能的解决方案

问题说明

你需要从DataFrame的job_description列中,提取每行文本里所有属于指定技能列表的内容,但当前代码仅返回首个匹配项,无法满足需求。

原代码问题点

  1. skill_list函数找到第一个匹配项就直接return,终止遍历,导致只返回单个技能
  2. 使用split()拆分文本后匹配,无法识别Qualitative research、User Persona这类多词组合的技能

修正方案

1. 修正技能列表

先修正拼写错误并统一字符串格式:

skills = ['Qualitative research', 'wireframing', 'figma', 'frame x', 'miro', 'mockflow', 'User Persona', 'coding', 'empathy', 'sketch',
          'communication', 'problem solving']

2. 编写全匹配提取函数

遍历技能列表,收集所有出现在当前文本中的技能:

def extract_all_skills(desc):
    matched = []
    # 不区分大小写匹配,若需严格大小写可移除lower()
    desc_lower = desc.lower()
    for skill in skills:
        if skill.lower() in desc_lower:
            matched.append(skill)
    return matched if matched else None  # 无匹配时返回None或空列表

3. 应用到DataFrame

用apply将函数作用于job_description列:

all_files['skills'] = all_files['job_description'].apply(extract_all_skills)

进阶:精确匹配(避免部分匹配)

如果要防止类似figma被误匹配为figmax,可以用正则做完整短语匹配:

import re

def extract_exact_skills(desc):
    matched = []
    for skill in skills:
        # 构建精确匹配的正则模式
        pattern = re.compile(r'\b' + re.escape(skill) + r'\b', re.IGNORECASE)
        if pattern.search(desc):
            matched.append(skill)
    return matched if matched else None

all_files['skills'] = all_files['job_description'].apply(extract_exact_skills)

最终效果

修改后skills列会包含该行所有匹配的技能,比如某行描述同时提到wireframing和figma,对应列值就是['wireframing', 'figma'],符合你的预期。

内容的提问来源于stack exchange,提问作者Kingsley Ihemere

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 05:10:31