无训练数据时,如何用Python+HuggingFace Transformers完成书籍预定义分类?
基于Python实现书籍预定义类别分类的解决方案
核心思路
无需依赖自定义标注数据,通过规则匹配+零样本预训练模型的组合就能低成本解决你的分类需求,既能快速落地,也能覆盖边缘场景。
具体实现方案
1. 规则匹配(快速落地)
直接利用标题、机构名称中的关键词匹配预定义类别,适合大部分明确的书籍:
- 农业类:匹配
agricultural、farm、农业经济协会等关键词 - 工程类:匹配
engineering、tech、机械等关键词 - 医疗类:匹配
medical、health、医院等关键词
示例代码:
def rule_based_classify(title, institution): # 统一转为小写避免大小写干扰 title_lower = title.lower() inst_lower = institution.lower() # 预定义各类别关键词 agri_keys = ['agricultural', 'farm', 'crop', 'livestock', 'agricultural economics association'] eng_keys = ['engineering', 'tech', 'mechanical', 'electrical', 'civil'] med_keys = ['medical', 'health', 'hospital', 'clinic', 'pharmaceutical'] if any(key in title_lower or key in inst_lower for key in agri_keys): return '农业' elif any(key in title_lower or key in inst_lower for key in eng_keys): return '工程' elif any(key in title_lower or key in inst_lower for key in med_keys): return '医疗' else: return '未分类' # 测试示例书籍 test_title = "Agriculturalpolicy and the decisions of agricultural producers as to income and investment" test_inst = "Western Agricultural Economics Association" print(rule_based_classify(test_title, test_inst)) # 输出:农业
2. 零样本预训练模型(覆盖边缘情况)
用HuggingFace的零样本分类模型,无需自定义标注数据,直接输入类别标签就能完成分类,适合关键词不明显的书籍:
示例代码:
from transformers import pipeline # 加载轻量零样本分类模型 classifier = pipeline("zero-shot-classification", model="cross-encoder/nli-MiniLM2-L6-H768") def zero_shot_classify(title, institution, candidate_labels): # 拼接标题和机构作为模型输入 input_text = f"书籍标题:{title},所属机构:{institution}" result = classifier(input_text, candidate_labels) # 返回得分最高的类别 return result['labels'][0] # 测试示例书籍 candidate_labels = ["农业", "工程", "医疗"] test_title = "Agriculturalpolicy and the decisions of agricultural producers as to income and investment" test_inst = "Western Agricultural Economics Association" print(zero_shot_classify(test_title, test_inst, candidate_labels)) # 输出:农业
3. 混合方案(兼顾速度与准确率)
先执行规则匹配,匹配成功直接返回结果;未匹配到的再调用零样本模型,平衡处理速度和分类覆盖范围。
优化建议
- 规则匹配可根据你的书籍库补充专业关键词,比如农业类新增
agronomy、horticulture等术语 - 零样本模型可按需更换,追求速度选轻量版,追求准确率选
facebook/bart-large-mnli
内容的提问来源于stack exchange,提问作者Yuri R
相关产品推荐
相关产品推荐

