Python中GPT-3大提示词使用及docx信息提取超限解决方案咨询
解决方案:处理超长docx内容的GPT-3信息提取问题
方法1:切换到支持更长上下文的模型
GPT-3基础模型(如text-davinci-003)仅支持4k tokens的上下文窗口,而GPT-3.5-turbo-16k(16k tokens)、GPT-4(8k/32k tokens)能直接容纳大部分docx文件的完整内容,无需拆分即可提取整体信息。
Python示例代码
import docx from openai import OpenAI client = OpenAI(api_key="你的API密钥") # 读取docx文件内容 def read_docx(file_path): doc = docx.Document(file_path) full_text = [] for para in doc.paragraphs: full_text.append(para.text) return "\n".join(full_text) docx_content = read_docx("目标文件路径.docx") # 调用大上下文模型提取信息 response = client.chat.completions.create( model="gpt-3.5-turbo-16k", messages=[ {"role": "system", "content": "请从以下文本中提取[具体信息类型,例如:所有项目名称、对应负责人及截止日期]"}, {"role": "user", "content": docx_content} ] ) print(response.choices[0].message.content)
方法2:分层摘要+整体提取
若必须使用GPT-3基础模型,可通过"分块生成保留上下文的摘要,再合并提取"的方式,既规避token限制,又保留整体信息关联:
- 将docx内容拆分为多个不超过GPT-3 token上限的片段;
- 对每个片段生成包含核心上下文的详细摘要(而非仅提取局部信息);
- 将所有摘要拼接后,用GPT-3基于合并文本提取整体信息。
Python示例代码
import docx from openai import OpenAI import tiktoken client = OpenAI(api_key="你的API密钥") tokenizer = tiktoken.get_encoding("cl100k_base") # 按token数拆分内容 def split_content_into_chunks(content, max_tokens=3000): tokens = tokenizer.encode(content) chunks = [] for i in range(0, len(tokens), max_tokens): chunk_tokens = tokens[i:i+max_tokens] chunks.append(tokenizer.decode(chunk_tokens)) return chunks # 生成单片段的上下文保留型摘要 def generate_chunk_summary(chunk): response = client.completions.create( model="text-davinci-003", prompt=f"请生成以下文本的详细摘要,完整保留关键上下文、核心事件及关联信息:\n{chunk}", max_tokens=500, temperature=0.1 ) return response.choices[0].text.strip() # 主执行流程 docx_content = read_docx("目标文件路径.docx") content_chunks = split_content_into_chunks(docx_content) chunk_summaries = [generate_chunk_summary(chunk) for chunk in content_chunks] merged_summary = "\n\n".join(chunk_summaries) # 基于合并摘要提取整体信息 final_response = client.completions.create( model="text-davinci-003", prompt=f"请从以下合并的文本摘要中提取[具体信息类型,例如:跨项目的依赖关系及整体进度状态]:\n{merged_summary}", max_tokens=1000, temperature=0.1 ) print(final_response.choices[0].text.strip())
方法3:结构化提取+迭代整合
如果目标是提取特定结构化信息(如实体、关键指标),可通过以下步骤实现整体整合:
- 定义清晰的提取模板(比如指定JSON字段);
- 分块提取每个片段中的对应字段;
- 调用GPT-3将所有分块结果整合,补全片段间的关联信息,去除重复内容。
示例思路代码
# 分块提取结构化信息 def extract_structured_info(chunk): response = client.completions.create( model="text-davinci-003", prompt=f"请从以下文本中提取项目信息,以JSON格式返回,包含字段:project_name, owner, deadline\n{chunk}", max_tokens=300, temperature=0 ) return response.choices[0].text.strip() # 整合所有提取结果 def integrate_extracted_results(all_extractions): response = client.completions.create( model="text-davinci-003", prompt=f"请整合以下所有项目信息,去除重复条目,补全跨项目的关联内容,最终以清晰的列表或JSON格式返回:\n{all_extractions}", max_tokens=1000, temperature=0.1 ) return response.choices[0].text.strip() # 执行流程 all_extractions = [extract_structured_info(chunk) for chunk in content_chunks] final_result = integrate_extracted_results("\n\n".join(all_extractions)) print(final_result)
内容的提问来源于stack exchange,提问作者TopTen1310
相关产品推荐
相关产品推荐

