连接Pinecone与OpenAI时出现MaxRetryError问题求助
Pinecone向量上传时MaxRetryError(SSLEOFError)问题排查与解决
我开发了一款简易应用,支持上传PDF文件、分割为文本块、生成Embedding并上传至Pinecone向量数据库。执行代码docsearch = Pinecone.from_texts([t.page_content for t in texts], embeddings, index_name=index_name)时触发如下MaxRetryError错误:
SSLEOFError Traceback (most recent call last) File /usr/lib/python3/dist-packages/urllib3/connectionpool.py:699, in HTTPConnectionPool.urlopen(self, method, url, body, headers, retries, redirect, assert_same_host, timeout, pool_timeout, release_conn, chunked, body_pos, **response_kw) 698 # Make the request on the httplib connection object. --> 699 httplib_response = self._make_request( 700 conn, 701 method, 702 url, 703 timeout=timeout_obj, 704 body=body, 705 headers=headers, 706 chunked=chunked, 707 ) 709 # If we're going to release the connection in ``finally:``, then 710 # the response doesn't need to know about the connection. Otherwise 711 # it will also try to release it and we'll have a double-release 712 # mess. File /usr/lib/python3/dist-packages/urllib3/connectionpool.py:394, in HTTPConnectionPool._make_request(self, conn, method, url, timeout, chunked, **httplib_request_kw) 393 else: --> 394 conn.request(method, url, **httplib_request_kw) 396 # We are swallowing BrokenPipeError (errno.EPIPE) since the server is 397 # legitimately able to close the connection after sending a valid response. 398 # With this behaviour, the received response is still readable. ... --> 574 raise MaxRetryError(_pool, url, error or ResponseError(cause)) 576 log.debug("Incremented Retry for (url='%s'): %r", url, new_retry) 578 return new_retry MaxRetryError: HTTPSConnectionPool(host='langchain2-e630e5d.svc.asia-northeast1-gcp.pinecone.io', port=443): Max retries exceeded with url: /vectors/upsert (Caused by SSLError(SSLEOFError(8, 'EOF occurred in violation of protocol (_ssl.c:2396)')))
完整代码如下:
from langchain.text_splitter import RecursiveCharacterTextSplitter # 加载数据 loader = UnstructuredPDFLoader("../data/field-guide-to-data-science.pdf") # loader = OnlinePDFLoader("https://wolfpaulus.com/wp-content/uploads/2017/05/field-guide-to-data-science.pdf") data = loader.load() print (f'You have {len(data)} document(s) in your data') print (f'There are {len(data[0].page_content)} characters in your document') # 输出: # You have 1 document(s) in your data # There are 176584 characters in your document # 将数据分割为更小的文本块 text_splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=0) texts = text_splitter.split_documents(data) print (f'Now you have {len(texts)} documents') # 输出: # Now you have 228 documents # 生成文档Embedding以准备语义搜索 from langchain.vectorstores import Chroma, Pinecone from langchain.embeddings.openai import OpenAIEmbeddings import pinecone OPENAI_API_KEY = '...' PINECONE_API_KEY = '...' PINECONE_API_ENV = 'us-east1-gcp' embeddings = OpenAIEmbeddings(openai_api_key=OPENAI_API_KEY) # 初始化Pinecone pinecone.init( api_key=PINECONE_API_KEY, # 在app.pinecone.io获取 environment=PINECONE_API_ENV # 控制台中API密钥旁查看 ) index_name = "langchain2" docsearch = Pinecone.from_texts([t.page_content for t in texts], embeddings, index_name=index_name)
问题排查与解决方案
核心问题:环境配置不匹配
错误信息中的Pinecone服务地址属于asia-northeast1-gcp环境,但代码里指定的PINECONE_API_ENV是us-east1-gcp,两者不匹配导致连接失败。登录Pinecone控制台,确认目标索引所在的环境,修改代码中的PINECONE_API_ENV为对应值(比如asia-northeast1-gcp)。网络连接验证
检查本地网络是否能访问Pinecone的目标服务地址,可执行如下命令测试:curl https://langchain2-e630e5d.svc.asia-northeast1-gcp.pinecone.io/vectors/upsert若无法连接,需排查防火墙规则、代理设置,确保网络能正常通达Pinecone的服务端口。
更新依赖版本
版本不兼容可能引发SSL连接问题,更新相关依赖到稳定版本:pip install --upgrade pinecone-client langchain openai分批次上传数据
一次性上传228个文本块可能导致请求过大,尝试分批次上传:batch_size = 50 for i in range(0, len(texts), batch_size): batch_texts = texts[i:i+batch_size] Pinecone.from_texts( [t.page_content for t in batch_texts], embeddings, index_name=index_name )临时绕过SSL验证(仅测试用)
若本地SSL证书存在问题,可临时关闭SSL验证(生产环境不建议使用):import ssl ssl._create_default_https_context = ssl._create_unverified_context
内容的提问来源于stack exchange,提问作者Nordic Guy
相关产品推荐
相关产品推荐

