You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

OpenAI分类功能已下线,如何替代实现数据分类?

问题描述

我编写了一段用于数据分类的代码:

result=openai.File.create(file=open("train.jsonl"),purpose="classifications")

执行后出现如下错误:

---------------------------------------------------------------------------
InvalidRequestError                       Traceback (most recent call last)
<ipython-input-34-c3332cf65668> in <cell line: 2>()
----> 1 result=openai.File.create(file=open("train.jsonl"), purpose="classifications")

3 frames
/usr/local/lib/python3.10/dist-packages/openai/api_requestor.py in _interpret_response_line(self, rbody,      rcode, rheaders, stream)
    681         stream_error = stream and "error" in resp.data
    682         if stream_error or not 200 <= rcode < 300:
--> 683             raise self.handle_error_response(
    684                 rbody, rcode, resp.data, rheaders, stream_error=stream_error
    685             )

InvalidRequestError: 'classifications' is not one of ['fine-tune'] - 'purpose'

已知OpenAI在2021年移除了classification、answers、search功能,请问不使用原分类功能的情况下,如何完成数据分类任务?

可行解决方案

1. 利用Chat Completions API做零样本/少样本分类

直接调用gpt-3.5-turbo或gpt-4模型,通过提示词引导模型完成分类,无需上传训练文件,适合中小规模分类场景。

零样本分类示例(无需训练数据)

import openai

openai.api_key = "你的API密钥"

def classify_text(text, categories):
    prompt = f"请将以下文本分类到指定类别中:{categories}。文本内容:{text}。只需返回对应的类别名称。"
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[
            {"role": "user", "content": prompt}
        ]
    )
    return response.choices[0].message['content'].strip()

# 使用示例
text_to_classify = "今天的股市指数上涨了2个百分点"
categories = ["财经", "体育", "娱乐", "科技"]
print(classify_text(text_to_classify, categories))

少样本分类示例(提供少量示例提升准确率)

如果分类边界模糊或类别特殊,提供少量标注示例能显著提升模型分类准确性:

import openai

openai.api_key = "你的API密钥"

def classify_text_with_examples(text, examples, categories):
    example_prompt = "\n".join([f"文本:{ex['text']},类别:{ex['category']}" for ex in examples])
    prompt = f"以下是分类示例:\n{example_prompt}\n请按照上述示例,将文本分类到指定类别:{categories}。待分类文本:{text}。只需返回类别名称。"
    response = openai.ChatCompletion.create(
        model="gpt-3.5-turbo",
        messages=[
            {"role": "user", "content": prompt}
        ]
    )
    return response.choices[0].message['content'].strip()

# 使用示例
examples = [
    {"text": "球队赢得了联赛冠军", "category": "体育"},
    {"text": "新发布的手机搭载了最新处理器", "category": "科技"}
]
text_to_classify = "歌手发布了新专辑,销量突破百万"
categories = ["财经", "体育", "娱乐", "科技"]
print(classify_text_with_examples(text_to_classify, examples, categories))

2. 微调OpenAI模型

如果有大量标注训练数据(通常建议至少几百条),可以微调gpt-3.5-turbo-instruct或babbage-002等模型,获得更高的分类效率和准确率,适合大规模或高要求的分类场景。步骤如下:

  1. 准备符合格式要求的训练数据(JSONL格式,每条数据包含prompt和completion字段),示例:
    {"prompt": "文本:今天的股市指数上涨了2个百分点\n类别:", "completion": " 财经"}
    {"prompt": "文本:球队赢得了联赛冠军\n类别:", "completion": " 体育"}
    
  2. 上传训练文件到OpenAI,注意purpose参数必须设为fine-tune:
    result = openai.File.create(file=open("train.jsonl"), purpose="fine-tune")
    
  3. 创建微调任务:
    openai.FineTuningJob.create(training_file=result.id, model="gpt-3.5-turbo-instruct")
    
  4. 等待微调任务完成后,使用生成的微调模型ID进行分类调用。

3. 结合Embedding API与传统分类器

先通过text-embedding-ada-002模型将文本转换为向量,再使用传统机器学习分类器(如逻辑回归、SVM)完成分类,适合需要本地部署或对成本敏感的场景:

  1. 生成文本向量:
    import openai
    
    openai.api_key = "你的API密钥"
    
    def get_embedding(text, model="text-embedding-ada-002"):
        text = text.replace("\n", " ")
        return openai.Embedding.create(input = [text], model=model)['data'][0]['embedding']
    
  2. 训练传统分类器(以逻辑回归为例):
    from sklearn.linear_model import LogisticRegression
    import numpy as np
    
    # 假设已准备好标注数据:texts(待训练文本列表)、labels(对应类别标签)
    embeddings = [get_embedding(text) for text in texts]
    X = np.array(embeddings)
    y = np.array(labels)
    
    # 训练模型
    clf = LogisticRegression()
    clf.fit(X, y)
    
    # 预测新文本
    new_text = "歌手发布了新专辑,销量突破百万"
    new_embedding = get_embedding(new_text)
    predicted_label = clf.predict([new_embedding])
    print(predicted_label)
    

内容的提问来源于stack exchange,提问作者Archa

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 03:05:35