You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

数据科学入门:如何用Python按关键词集合为DataFrame生成分类编码

解答

1. 现有基础代码的简化方案

你目前写的测试代码存在冗余,循环逻辑没有用到迭代的关键词变量,只是单纯重复输出E,直接保留关键词集合和对应编码的映射定义即可:

# 医疗类关键词集合
healthcare_keywords = {'healthcare', 'health care', 'health', 'hospital', 'medical'}
# 医疗类对应NTEE编码
healthcare_code = 'E'

2. DataFrame列批量赋值实现

不需要手动写循环遍历,用pandas内置的向量化操作即可完成,运行效率远高于自定义循环,代码如下:

import pandas as pd

# 生成关键词匹配正则,用|连接所有关键词表示任意匹配
match_pattern = '|'.join(healthcare_keywords)
# 符合匹配规则的行B列赋值为E,不符合的默认留空,可按需修改默认值
df['B'] = df['A'].str.contains(
    match_pattern,
    case=False, # 忽略大小写匹配
    na=False # 空值默认判定为不匹配,避免报错
).map({True: healthcare_code, False: ''})

如果后续需要扩展多分类规则,可以用numpy.where做多层条件判断,示例如下:

import numpy as np

# 示例:教育类关键词和对应编码
education_keywords = {'school', 'education', 'university', 'college'}
education_code = 'B'
edu_pattern = '|'.join(education_keywords)

# 多层条件赋值
df['B'] = np.where(
    df['A'].str.contains(match_pattern, case=False, na=False), healthcare_code,
    np.where(df['A'].str.contains(edu_pattern, case=False, na=False), education_code,
    '')
)

内容的提问来源于stack exchange,提问作者Dakotah Thunder Wilson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 11:27:05