You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何判断子主题归属后逐个插入至pandas DataFrame对应列?

子主题分类到DataFrame对应列的实现方案

问题概述

现有一个以通用主题为表头的DataFrame:

["General Pharmacology, Toxicology and Pharmaceutics", "General Medicine", "General Biochemistry, Genetics and Molecular Biology", "General Chemical Engineering", "General Dentistry", "Others"]

以及一个子主题列表,需要将每个子主题判定所属通用主题后,插入到DataFrame的对应列中。核心疑问是:能否通过遍历子主题来选择插入的目标列?

回答

完全可以通过遍历子主题实现该需求,核心是先建立子主题与通用主题的映射规则,再遍历每个子主题匹配目标列完成插入。以下是基于Python pandas的具体实现步骤:

1. 初始化空DataFrame

先创建符合通用主题表头的空DataFrame:

import pandas as pd

# 定义通用主题表头
general_topics = [
    "General Pharmacology, Toxicology and Pharmaceutics",
    "General Medicine",
    "General Biochemistry, Genetics and Molecular Biology",
    "General Chemical Engineering",
    "General Dentistry",
    "Others"
]

# 创建空DataFrame
df = pd.DataFrame(columns=general_topics)

2. 定义主题映射规则

根据子主题的领域属性,建立子主题到通用主题的映射字典(可根据实际需求调整):

topic_mapping = {
    # 药理、毒理与药剂学类
    "Health, Toxicology and Mutagenesis": "General Pharmacology, Toxicology and Pharmaceutics",
    "Chemical Health and Safety": "General Pharmacology, Toxicology and Pharmaceutics",
    # 综合医学类
    "Maternity and Midwifery": "General Medicine",
    "Emergency": "General Medicine",
    "Transplantation": "General Medicine",
    "Medical–Surgical": "General Medicine",
    "Biological Psychiatry": "General Medicine",
    "Dermatology": "General Medicine",
    "Medicine (miscellaneous)": "General Medicine",
    "Anesthesiology and Pain Medicine": "General Medicine",
    "Reproductive Medicine": "General Medicine",
    "Speech and Hearing": "General Medicine",
    "Physiology": "General Medicine",
    # 生物化学、遗传与分子生物学类
    "Biochemistry, Genetics and Molecular Biology (miscellaneous)": "General Biochemistry, Genetics and Molecular Biology",
    # 化学工程类
    "Colloid and Surface Chemistry": "General Chemical Engineering",
    "Chemistry (miscellaneous)": "General Chemical Engineering",
    "Inorganic Chemistry": "General Chemical Engineering",
    "Ceramics and Composites": "General Chemical Engineering",
    "Spectroscopy": "General Chemical Engineering",
    # 其余子主题默认归为Others
    "Management Science and Operations Research": "Others",
    "Economics and Econometrics": "Others",
    "Discrete Mathematics and Combinatorics": "Others",
    "Communication": "Others",
    "Psychology (miscellaneous)": "Others",
    "Urban Studies": "Others",
    "Geophysics": "Others",
    "Geology": "Others",
    "Neuropsychology and Physiological Psychology": "Others",
    "Space and Planetary Science": "Others",
    "Marketing": "Others",
    "Experimental and Cognitive Psychology": "Others",
    "Condensed Matter Physics": "Others",
    "Architecture ": "Others",
    "Small Animals": "Others",
    "Religious studies": "Others",
    "Insect Science": "Others",
    "Multidisciplinary": "Others",
    "Health Policy": "Others",
    "Atmospheric Science": "Others"
}

3. 遍历子主题并插入对应列

遍历每个子主题,通过映射找到目标列,将子主题追加到该列:

# 子主题列表
sub_topics = [
    'Spectroscopy', 'Maternity and Midwifery', 'Emergency',
    'Management Science and Operations Research', 'Economics and Econometrics',
    'Transplantation', 'Discrete Mathematics and Combinatorics', 'Communication',
    'Colloid and Surface Chemistry', 'Psychology (miscellaneous)',
    'Health, Toxicology and Mutagenesis', 'Medical–Surgical', 'Urban Studies',
    'Geophysics', 'Geology', 'Neuropsychology and Physiological Psychology',
    'Space and Planetary Science', 'Chemistry (miscellaneous)', 'Marketing',
    'Inorganic Chemistry', 'Experimental and Cognitive Psychology',
    'Condensed Matter Physics', 'Architecture ', 'Chemical Health and Safety',
    'Biological Psychiatry', 'Dermatology', 'Medicine (miscellaneous)',
    'Biochemistry, Genetics and Molecular Biology (miscellaneous)', 'Speech and Hearing',
    'Anesthesiology and Pain Medicine', 'Reproductive Medicine', 'Small Animals',
    'Religious studies', 'Insect Science', 'Ceramics and Composites', 'Multidisciplinary',
    'Health Policy', 'Atmospheric Science', 'Physiology'
]

# 遍历子主题,插入对应列
for sub_topic in sub_topics:
    # 获取目标列,默认归为Others
    target_col = topic_mapping.get(sub_topic, "Others")
    # 追加到目标列末尾
    df.loc[len(df), target_col] = sub_topic

4. 优化DataFrame(可选)

将空值填充为空字符串,让结果更整洁:

df = df.fillna("")

说明

这种遍历方式逻辑直观,便于后续维护和调整映射规则。如果子主题数量极大,可以考虑批量处理优化效率,但对于常规规模的子主题列表,遍历的方式完全够用。

内容的提问来源于stack exchange,提问作者Marlon Teixeira

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.16 23:30:57