You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

无训练数据时,如何用Python+HuggingFace Transformers完成书籍预定义分类?

基于Python实现书籍预定义类别分类的解决方案

核心思路

无需依赖自定义标注数据,通过规则匹配+零样本预训练模型的组合就能低成本解决你的分类需求,既能快速落地,也能覆盖边缘场景。

具体实现方案

1. 规则匹配(快速落地)

直接利用标题、机构名称中的关键词匹配预定义类别,适合大部分明确的书籍:

  • 农业类:匹配agricultural、farm、农业经济协会等关键词
  • 工程类:匹配engineering、tech、机械等关键词
  • 医疗类:匹配medical、health、医院等关键词

示例代码:

def rule_based_classify(title, institution):
    # 统一转为小写避免大小写干扰
    title_lower = title.lower()
    inst_lower = institution.lower()
    
    # 预定义各类别关键词
    agri_keys = ['agricultural', 'farm', 'crop', 'livestock', 'agricultural economics association']
    eng_keys = ['engineering', 'tech', 'mechanical', 'electrical', 'civil']
    med_keys = ['medical', 'health', 'hospital', 'clinic', 'pharmaceutical']
    
    if any(key in title_lower or key in inst_lower for key in agri_keys):
        return '农业'
    elif any(key in title_lower or key in inst_lower for key in eng_keys):
        return '工程'
    elif any(key in title_lower or key in inst_lower for key in med_keys):
        return '医疗'
    else:
        return '未分类'

# 测试示例书籍
test_title = "Agriculturalpolicy and the decisions of agricultural producers as to income and investment"
test_inst = "Western Agricultural Economics Association"
print(rule_based_classify(test_title, test_inst))  # 输出:农业

2. 零样本预训练模型(覆盖边缘情况)

用HuggingFace的零样本分类模型,无需自定义标注数据,直接输入类别标签就能完成分类,适合关键词不明显的书籍:

示例代码:

from transformers import pipeline

# 加载轻量零样本分类模型
classifier = pipeline("zero-shot-classification", model="cross-encoder/nli-MiniLM2-L6-H768")

def zero_shot_classify(title, institution, candidate_labels):
    # 拼接标题和机构作为模型输入
    input_text = f"书籍标题:{title},所属机构:{institution}"
    result = classifier(input_text, candidate_labels)
    # 返回得分最高的类别
    return result['labels'][0]

# 测试示例书籍
candidate_labels = ["农业", "工程", "医疗"]
test_title = "Agriculturalpolicy and the decisions of agricultural producers as to income and investment"
test_inst = "Western Agricultural Economics Association"
print(zero_shot_classify(test_title, test_inst, candidate_labels))  # 输出:农业

3. 混合方案(兼顾速度与准确率)

先执行规则匹配,匹配成功直接返回结果;未匹配到的再调用零样本模型,平衡处理速度和分类覆盖范围。

优化建议

  • 规则匹配可根据你的书籍库补充专业关键词,比如农业类新增agronomy、horticulture等术语
  • 零样本模型可按需更换,追求速度选轻量版,追求准确率选facebook/bart-large-mnli

内容的提问来源于stack exchange,提问作者Yuri R

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.12 17:42:40