You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用OpenAI GPT-3 API生成更长文本?技术求助

解决GPT-3 API生成文本过短的问题

问题核心原因

你的输出远短于预期,主要由三个问题导致:

  • 惩罚参数过高:frequency_penalty=1和presence_penalty=1会强烈抑制模型重复内容、拓展新主题,直接导致生成提前终止。
  • prompt引导性不足:简短的"How to choose a student loan"没有明确要求模型展开细节,只会触发概括性回答。
  • 对n参数的误解:n=10是生成10个独立的候选结果,而非连续拼接的长文本,所以取第一个结果依然是短内容。

具体解决方案

1. 调整惩罚参数

将frequency_penalty和presence_penalty降至0-0.3区间,既避免内容过度重复,又不会限制模型展开:

frequency_penalty=0.2,
presence_penalty=0.1

2. 优化prompt明确内容要求

修改prompt,明确指定生成长度、覆盖的核心要点,引导模型输出详细内容:

prompt="Write a comprehensive guide (around 1000 tokens) on choosing a student loan. Cover key aspects: fixed vs variable interest rates, federal vs private loan differences, repayment plan options, eligibility criteria, how to compare loan offers, and tips to minimize long-term costs."

3. 设置对应长度的max_tokens

要生成约1000token的文本,直接将max_tokens设为1000(注意:模型总上下文token数不能超过4096,需确保prompt的token数+1000不超出该限制):

max_tokens=1000

4. 续写实现超长文本(可选)

如果单次生成仍达不到目标长度,可以采用续写逻辑:将上一次的生成内容拼接回prompt,重复调用API直到满足长度要求。示例代码:

import os
import openai

openai.api_key = os.getenv("OPENAI_API_KEY")

def generate_long_text(base_prompt, target_tokens=1000):
    current_content = ""
    remaining_tokens = target_tokens
    while remaining_tokens > 0:
        generate_num = min(remaining_tokens, 1000)
        response = openai.Completion.create(
            model="text-davinci-002",
            prompt=base_prompt + current_content,
            temperature=0.6,
            max_tokens=generate_num,
            top_p=1,
            frequency_penalty=0.2,
            presence_penalty=0.1,
            n=1
        )
        new_segment = response['choices'][0]['text'].strip()
        if not new_segment:
            break
        current_content += " " + new_segment
        # 粗略估算新增token数(如需精确计算可使用tiktoken库)
        remaining_tokens -= len(new_segment.split()) * 1.3
    return current_content

# 使用示例
base_prompt = "Write a comprehensive guide on choosing a student loan. Cover key aspects: fixed vs variable interest rates, federal vs private loan differences, repayment plan options, eligibility criteria, how to compare loan offers, and tips to minimize long-term costs."
long_result = generate_long_text(base_prompt, target_tokens=1000)
print(long_result)

长度验证

如果需要精确统计生成的token数,可以安装tiktoken库(pip install tiktoken),通过以下代码计算:

import tiktoken
encoding = tiktoken.get_encoding("cl100k_base")
token_count = len(encoding.encode(long_result))
print(f"生成文本的token数:{token_count}")

内容的提问来源于stack exchange,提问作者Aydin Abiar

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 23:33:22