PDF上传读取失败:AttributeError: 'bytes'对象无seek属性求助
解决多PDF读取应用的
AttributeError: 'bytes' object has no attribute 'seek'问题 错误信息
AttributeError: 'bytes' object has no attribute 'seek'
回溯信息
File "c:\Users\XANDER\OneDrive\Desktop\Python Projects\.venv\Lib\site-packages\streamlit\runtime\scriptrunner\script_runner.py", line 600, in _run_script exec(code, module.__dict__) File "C:\Users\XANDER\OneDrive\Desktop\Python Projects\app.py", line 89, in <module> main() File "C:\Users\XANDER\OneDrive\Desktop\Python Projects\app.py", line 83, in main raw_text = get_pdf_text(pdf_docs) ^^^^^^^^^^^^^^^^^^^^^^ File "C:\Users\XANDER\OneDrive\Desktop\Python Projects\app.py", line 21, in get_pdf_text pdf_reader=PdfReader(pdf) ^^^^^^^^^^^^^^ File "c:\Users\XANDER\OneDrive\Desktop\Python Projects\.venv\Lib\site-packages\pypdf\_reader.py", line 127, in __init__ self.read(stream) File "c:\Users\XANDER\OneDrive\Desktop\Python Projects\.venv\Lib\site-packages\pypdf\_reader.py", line 538, in read self._basic_validation(stream) File "c:\Users\XANDER\OneDrive\Desktop\Python Projects\.venv\Lib\site-packages\pypdf\_reader.py", line 597, in _basic_validation stream.seek(0, os.SEEK_SET) ^^^^^^^^^^^
问题原因
- 核心问题:Streamlit的
st.file_uploader默认返回单个UploadedFile对象(上传单个文件时),但你的get_pdf_text函数预期接收文件列表并遍历。遍历单个UploadedFile对象时,会将其解析为字节流,导致pdf变量变成bytes类型,而PdfReader需要可seek的文件对象,因此触发报错。 - 代码还存在两个次要问题:
- 对话模板里的变量
{queestion}拼写错误,应为{question} user_input函数中仅引用了get_conversational_chain函数对象,未实际调用
- 对话模板里的变量
修复方案
1. 修改文件上传组件
在main函数的侧边栏中,给st.file_uploader添加accept_multiple_files=True参数,确保无论上传单个还是多个文件,都返回列表:
pdf_docs = st.file_uploader("Upload your PDF Files and Click on Submit & Process", accept_multiple_files=True)
2. 增加空值判断(可选但推荐)
在get_pdf_text函数开头添加判断,避免空列表传入时出错:
def get_pdf_text(pdf_docs): text="" if not pdf_docs: return text for pdf in pdf_docs: pdf_reader=PdfReader(pdf) for page in pdf_reader.pages: text+=page.extract_text() return text
3. 修正拼写错误
在get_conversational_chain函数的prompt模板中,将{queestion}改为{question}:
prompt_template=""" Answer the question as detailed as possible from the provided context, make sure to provide all the details, if the answer is not in the provided context just say, "answer is not available in the context", don't provide the wrong answer \n\n Context:\n {context}?\n Question:\n{question}\n Answer: """
4. 修复对话链调用
在user_input函数中,调用get_conversational_chain函数:
chain = get_conversational_chain()
修改后的完整代码
import streamlit as st from pypdf import PdfReader from langchain.text_splitter import RecursiveCharacterTextSplitter import os from langchain_google_genai import GoogleGenerativeAIEmbeddings import google.generativeai as genai from langchain_community.vectorstores import FAISS from langchain_google_genai import ChatGoogleGenerativeAI from langchain.chains.question_answering import load_qa_chain from langchain.prompts import PromptTemplate from dotenv import load_dotenv load_dotenv() genai.configure(api_key=os.getenv("GOOGLE_API_KEY")) def get_pdf_text(pdf_docs): text="" if not pdf_docs: return text for pdf in pdf_docs: pdf_reader=PdfReader(pdf) for page in pdf_reader.pages: text+=page.extract_text() return text def get_text_chunks(text): text_splitter=RecursiveCharacterTextSplitter(chunk_size=10000, chunk_overlap=1000) chunks=text_splitter.split_text(text) return chunks def get_vector_store(text_chunks): embeddings=GoogleGenerativeAIEmbeddings(model="models/embedding-001") vector_store=FAISS.from_texts(text_chunks, embedding=embeddings) vector_store.save_local("faiss_index") def get_conversational_chain(): prompt_template=""" Answer the question as detailed as possible from the provided context, make sure to provide all the details, if the answer is not in the provided context just say, "answer is not available in the context", don't provide the wrong answer \n\n Context:\n {context}?\n Question:\n{question}\n Answer: """ model=ChatGoogleGenerativeAI(model="gemini-pro", temperature=0.3) prompt=PromptTemplate(template=prompt_template, input_variables=["context", "question"]) chain=load_qa_chain(model, chain_type="stuff", prompt=prompt) return chain def user_input(user_question): embeddings=GoogleGenerativeAIEmbeddings(model="models/embedding-001") new_db = FAISS.load_local("faiss_index", embeddings, allow_dangerous_deserialization=True) docs = new_db.similarity_search(user_question) chain = get_conversational_chain() response = chain( {"input_documents":docs, "question": user_question} , return_only_outputs=True ) print(response) st.write("Reply: ", response["output_text"]) def main(): st.set_page_config("Chat With Multiple PDFs") st.header("Chat with Multiple PDFs using Gemini 💁♀️") user_question = st.text_input("Ask a Question from the PDF Files") if user_question: user_input(user_question) with st.sidebar: st.title("Menu:") pdf_docs = st.file_uploader("Upload your PDF Files and Click on Submit & Process", accept_multiple_files=True) if st.button("Submit & Process"): with st.spinner("Processing..."): raw_text = get_pdf_text(pdf_docs) text_chunks = get_text_chunks(raw_text) get_vector_store(text_chunks) st.success("Done") if __name__ == "__main__": main()
注:额外添加了allow_dangerous_deserialization=True到FAISS.load_local,避免新版本FAISS的序列化警告
内容的提问来源于stack exchange,提问作者aria obscura
相关产品推荐
相关产品推荐

