如何使用GPT-3进行文本分类?含TensorFlow、Keras迁移学习入门方法
嘿,这个问题我刚好有不少实践经验,咱们一步步来拆解:
GPT-3做文本分类主要有两种思路,完全取决于你手头的数据集大小:
1. 零样本/少样本分类(小数据集首选)
这是GPT-3最省心的用法,不用训练,靠精心设计的prompt就能完成分类。核心是给模型清晰的任务指令,甚至可以给几个示例帮它快速理解分类规则。
举个情感分类的实际例子,你的prompt可以这么写:
请判断以下文本的情感类型,可选类别为:正面、负面、中性。
示例1:文本:这部电影的剧情太精彩了! 情感:正面
示例2:文本:今天的交通堵了整整一小时,糟透了! 情感:负面
文本:明天会下雨。 情感:
然后调用OpenAI的Completion API就能得到结果。用Python代码实现的话大概是这样:
import openai openai.api_key = "你的API密钥" def classify_text(text): prompt = """请判断以下文本的情感类型,可选类别为:正面、负面、中性。 示例1:文本:这部电影的剧情太精彩了! 情感:正面 示例2:文本:今天的交通堵了整整一小时,糟透了! 情感:负面 文本:{} 情感:""".format(text) response = openai.Completion.create( model="text-davinci-003", prompt=prompt, temperature=0, # 温度设为0让结果更稳定,避免随机波动 max_tokens=10 ) return response.choices[0].text.strip() # 测试一下 print(classify_text("这家餐厅的服务真周到"))
2. 微调GPT-3(大数据集提升效果)
如果你的数据集有几百上千条标注样本,微调GPT-3能让分类效果更稳定、准确率更高。步骤大概是:
- 准备数据集:把样本转换成JSONL格式,每条数据包含
prompt(输入文本)和completion(分类结果),比如:{"prompt": "文本:这家餐厅的服务真周到\n情感:", "completion": "正面"} {"prompt": "文本:今天的外卖迟到了半小时\n情感:", "completion": "负面"} - 上传数据集到OpenAI:用
openai api files.create命令上传JSONL文件 - 启动微调任务:运行
openai api fine_tunes.create -t 你的数据集文件名 -m davinci - 用微调后的模型做分类:调用Completion API时指定微调后的模型ID即可
完全可行!GPT-3在海量通用文本上预训练过,已经学到了极强的语言表示能力,我们可以把它作为特征提取器,将其输出的文本嵌入(embedding)迁移到特定的文本分类任务中,或者在微调GPT-3的基础上,进一步适配你的目标任务。
这种方式的优势很明显:不用从头训练大模型,借助GPT-3的预训练特征,能在小数据集上快速得到不错的效果。
这里我推荐「GPT-3嵌入 + Keras分类器」的方案,上手最快,效果也稳定:
步骤1:获取GPT-3的文本嵌入
首先,调用OpenAI的Embedding API,把所有训练和测试文本转换成1536维的嵌入向量(text-embedding-ada-002模型的输出维度)。代码示例:
import openai import numpy as np openai.api_key = "你的API密钥" def get_embedding(text): response = openai.Embedding.create( input=text, model="text-embedding-ada-002" ) return response['data'][0]['embedding'] # 假设你的训练数据是texts和labels texts = ["这家餐厅的服务真周到", "今天的外卖迟到了半小时", ...] labels = ["正面", "负面", ...] # 转换所有文本为嵌入向量 X = np.array([get_embedding(text) for text in texts]) # 标签编码(转成整数格式) from sklearn.preprocessing import LabelEncoder le = LabelEncoder() y = le.fit_transform(labels) # 如果是多分类任务,转成one-hot编码 from tensorflow.keras.utils import to_categorical y = to_categorical(y)
步骤2:构建Keras分类模型
用TensorFlow/Keras搭建一个简单的全连接模型,输入是GPT-3的嵌入向量,输出是分类结果:
from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense, Dropout model = Sequential([ Dense(256, activation='relu', input_shape=(1536,)), Dropout(0.3), # 防止过拟合 Dense(128, activation='relu'), Dropout(0.3), Dense(len(le.classes_), activation='softmax') # 输出层维度等于类别数 ]) model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])
步骤3:训练和评估模型
划分训练集和测试集,然后训练模型:
from sklearn.model_selection import train_test_split X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) # 开始训练 history = model.fit(X_train, y_train, epochs=10, batch_size=8, validation_split=0.1) # 评估模型效果 test_loss, test_acc = model.evaluate(X_test, y_test) print(f"测试准确率:{test_acc:.2f}")
步骤4:用模型做预测
拿到新文本后,先转成嵌入向量,再输入模型预测:
def predict_sentiment(text): embedding = get_embedding(text) embedding = np.expand_dims(embedding, axis=0) # 增加batch维度,适配模型输入要求 pred = model.predict(embedding)[0] return le.inverse_transform([np.argmax(pred)])[0] # 测试一下 print(predict_sentiment("这家店的菜太咸了,不会再来了"))
内容的提问来源于stack exchange,提问作者anveshtummala

