You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中GPT-3大提示词使用及docx信息提取超限解决方案咨询

解决方案:处理超长docx内容的GPT-3信息提取问题

方法1:切换到支持更长上下文的模型

GPT-3基础模型(如text-davinci-003)仅支持4k tokens的上下文窗口,而GPT-3.5-turbo-16k(16k tokens)、GPT-4(8k/32k tokens)能直接容纳大部分docx文件的完整内容,无需拆分即可提取整体信息。

Python示例代码

import docx
from openai import OpenAI

client = OpenAI(api_key="你的API密钥")

# 读取docx文件内容
def read_docx(file_path):
    doc = docx.Document(file_path)
    full_text = []
    for para in doc.paragraphs:
        full_text.append(para.text)
    return "\n".join(full_text)

docx_content = read_docx("目标文件路径.docx")

# 调用大上下文模型提取信息
response = client.chat.completions.create(
    model="gpt-3.5-turbo-16k",
    messages=[
        {"role": "system", "content": "请从以下文本中提取[具体信息类型,例如:所有项目名称、对应负责人及截止日期]"},
        {"role": "user", "content": docx_content}
    ]
)

print(response.choices[0].message.content)

方法2:分层摘要+整体提取

若必须使用GPT-3基础模型,可通过"分块生成保留上下文的摘要,再合并提取"的方式,既规避token限制,又保留整体信息关联:

  1. 将docx内容拆分为多个不超过GPT-3 token上限的片段;
  2. 对每个片段生成包含核心上下文的详细摘要(而非仅提取局部信息);
  3. 将所有摘要拼接后,用GPT-3基于合并文本提取整体信息。

Python示例代码

import docx
from openai import OpenAI
import tiktoken

client = OpenAI(api_key="你的API密钥")
tokenizer = tiktoken.get_encoding("cl100k_base")

# 按token数拆分内容
def split_content_into_chunks(content, max_tokens=3000):
    tokens = tokenizer.encode(content)
    chunks = []
    for i in range(0, len(tokens), max_tokens):
        chunk_tokens = tokens[i:i+max_tokens]
        chunks.append(tokenizer.decode(chunk_tokens))
    return chunks

# 生成单片段的上下文保留型摘要
def generate_chunk_summary(chunk):
    response = client.completions.create(
        model="text-davinci-003",
        prompt=f"请生成以下文本的详细摘要,完整保留关键上下文、核心事件及关联信息:\n{chunk}",
        max_tokens=500,
        temperature=0.1
    )
    return response.choices[0].text.strip()

# 主执行流程
docx_content = read_docx("目标文件路径.docx")
content_chunks = split_content_into_chunks(docx_content)
chunk_summaries = [generate_chunk_summary(chunk) for chunk in content_chunks]
merged_summary = "\n\n".join(chunk_summaries)

# 基于合并摘要提取整体信息
final_response = client.completions.create(
    model="text-davinci-003",
    prompt=f"请从以下合并的文本摘要中提取[具体信息类型,例如:跨项目的依赖关系及整体进度状态]:\n{merged_summary}",
    max_tokens=1000,
    temperature=0.1
)

print(final_response.choices[0].text.strip())

方法3:结构化提取+迭代整合

如果目标是提取特定结构化信息(如实体、关键指标),可通过以下步骤实现整体整合:

  1. 定义清晰的提取模板(比如指定JSON字段);
  2. 分块提取每个片段中的对应字段;
  3. 调用GPT-3将所有分块结果整合,补全片段间的关联信息,去除重复内容。

示例思路代码

# 分块提取结构化信息
def extract_structured_info(chunk):
    response = client.completions.create(
        model="text-davinci-003",
        prompt=f"请从以下文本中提取项目信息,以JSON格式返回,包含字段:project_name, owner, deadline\n{chunk}",
        max_tokens=300,
        temperature=0
    )
    return response.choices[0].text.strip()

# 整合所有提取结果
def integrate_extracted_results(all_extractions):
    response = client.completions.create(
        model="text-davinci-003",
        prompt=f"请整合以下所有项目信息,去除重复条目,补全跨项目的关联内容,最终以清晰的列表或JSON格式返回:\n{all_extractions}",
        max_tokens=1000,
        temperature=0.1
    )
    return response.choices[0].text.strip()

# 执行流程
all_extractions = [extract_structured_info(chunk) for chunk in content_chunks]
final_result = integrate_extracted_results("\n\n".join(all_extractions))
print(final_result)

内容的提问来源于stack exchange,提问作者TopTen1310

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.22 15:35:25