OpenAI分类功能已下线,如何替代实现数据分类?
问题描述
我编写了一段用于数据分类的代码:
result=openai.File.create(file=open("train.jsonl"),purpose="classifications")
执行后出现如下错误:
--------------------------------------------------------------------------- InvalidRequestError Traceback (most recent call last) <ipython-input-34-c3332cf65668> in <cell line: 2>() ----> 1 result=openai.File.create(file=open("train.jsonl"), purpose="classifications") 3 frames /usr/local/lib/python3.10/dist-packages/openai/api_requestor.py in _interpret_response_line(self, rbody, rcode, rheaders, stream) 681 stream_error = stream and "error" in resp.data 682 if stream_error or not 200 <= rcode < 300: --> 683 raise self.handle_error_response( 684 rbody, rcode, resp.data, rheaders, stream_error=stream_error 685 ) InvalidRequestError: 'classifications' is not one of ['fine-tune'] - 'purpose'
已知OpenAI在2021年移除了classification、answers、search功能,请问不使用原分类功能的情况下,如何完成数据分类任务?
可行解决方案
1. 利用Chat Completions API做零样本/少样本分类
直接调用gpt-3.5-turbo或gpt-4模型,通过提示词引导模型完成分类,无需上传训练文件,适合中小规模分类场景。
零样本分类示例(无需训练数据)
import openai openai.api_key = "你的API密钥" def classify_text(text, categories): prompt = f"请将以下文本分类到指定类别中:{categories}。文本内容:{text}。只需返回对应的类别名称。" response = openai.ChatCompletion.create( model="gpt-3.5-turbo", messages=[ {"role": "user", "content": prompt} ] ) return response.choices[0].message['content'].strip() # 使用示例 text_to_classify = "今天的股市指数上涨了2个百分点" categories = ["财经", "体育", "娱乐", "科技"] print(classify_text(text_to_classify, categories))
少样本分类示例(提供少量示例提升准确率)
如果分类边界模糊或类别特殊,提供少量标注示例能显著提升模型分类准确性:
import openai openai.api_key = "你的API密钥" def classify_text_with_examples(text, examples, categories): example_prompt = "\n".join([f"文本:{ex['text']},类别:{ex['category']}" for ex in examples]) prompt = f"以下是分类示例:\n{example_prompt}\n请按照上述示例,将文本分类到指定类别:{categories}。待分类文本:{text}。只需返回类别名称。" response = openai.ChatCompletion.create( model="gpt-3.5-turbo", messages=[ {"role": "user", "content": prompt} ] ) return response.choices[0].message['content'].strip() # 使用示例 examples = [ {"text": "球队赢得了联赛冠军", "category": "体育"}, {"text": "新发布的手机搭载了最新处理器", "category": "科技"} ] text_to_classify = "歌手发布了新专辑,销量突破百万" categories = ["财经", "体育", "娱乐", "科技"] print(classify_text_with_examples(text_to_classify, examples, categories))
2. 微调OpenAI模型
如果有大量标注训练数据(通常建议至少几百条),可以微调gpt-3.5-turbo-instruct或babbage-002等模型,获得更高的分类效率和准确率,适合大规模或高要求的分类场景。步骤如下:
- 准备符合格式要求的训练数据(JSONL格式,每条数据包含
prompt和completion字段),示例:{"prompt": "文本:今天的股市指数上涨了2个百分点\n类别:", "completion": " 财经"} {"prompt": "文本:球队赢得了联赛冠军\n类别:", "completion": " 体育"} - 上传训练文件到OpenAI,注意
purpose参数必须设为fine-tune:result = openai.File.create(file=open("train.jsonl"), purpose="fine-tune") - 创建微调任务:
openai.FineTuningJob.create(training_file=result.id, model="gpt-3.5-turbo-instruct") - 等待微调任务完成后,使用生成的微调模型ID进行分类调用。
3. 结合Embedding API与传统分类器
先通过text-embedding-ada-002模型将文本转换为向量,再使用传统机器学习分类器(如逻辑回归、SVM)完成分类,适合需要本地部署或对成本敏感的场景:
- 生成文本向量:
import openai openai.api_key = "你的API密钥" def get_embedding(text, model="text-embedding-ada-002"): text = text.replace("\n", " ") return openai.Embedding.create(input = [text], model=model)['data'][0]['embedding'] - 训练传统分类器(以逻辑回归为例):
from sklearn.linear_model import LogisticRegression import numpy as np # 假设已准备好标注数据:texts(待训练文本列表)、labels(对应类别标签) embeddings = [get_embedding(text) for text in texts] X = np.array(embeddings) y = np.array(labels) # 训练模型 clf = LogisticRegression() clf.fit(X, y) # 预测新文本 new_text = "歌手发布了新专辑,销量突破百万" new_embedding = get_embedding(new_text) predicted_label = clf.predict([new_embedding]) print(predicted_label)
内容的提问来源于stack exchange,提问作者Archa
相关产品推荐
相关产品推荐

