Azure OpenAI Embeddings调用Azure AI Search时Tiktoken连接超时问题求助
问题描述
使用Azure AI Search实例搭配text-embedding-ada-002嵌入模型,通过langchain_openai的AzureOpenAIEmbeddings调用嵌入函数无异常:
self.model = AzureOpenAIEmbeddings(model=self.embedding_deployment, azure_endpoint=self.endpoint, openai_api_key = self.api_key, openai_api_version="2024-02-01")
但使用langchain.community.vectorstores的AzureSearch创建索引时出现连接超时错误:
self.search_model = AzureSearch(azure_search_endpoint=self.azure_search_endpoint, azure_search_key=self.search_api_key, index_name=self.index_name, embedding_function=self.model.embed_query, ) #连接错误
错误详情:
Exception has occurred: ConnectTimeout HTTPSConnectionPool(host='openaipublic.blob.core.windows.net', port=443): Max retries exceeded with url: /encodings/cl100k_base.tiktoken (Caused by ConnectTimeoutError(<urllib3.connection.HTTPSConnection object at 0x0000020FFF8A1010>, 'Connection to openaipublic.blob.core.windows.net timed out. (connect timeout=None)')) KeyError: 'Could not automatically map <embedding_deployment_name> to a tokeniser. Please use `tiktoken.get_encoding` to explicitly get the tokeniser you expect.' During handling of the above exception, another exception occurred: TimeoutError: [WinError 10060] A connection attempt failed because the connected party did not properly respond after a period of time, or established connection failed because connected host has failed to respond During handling of the above exception, another exception occurred: urllib3.exceptions.ConnectTimeoutError: (<urllib3.connection.HTTPSConnection object at 0x0000020FFF8A1010>, 'Connection to openaipublic.blob.core.windows.net timed out. (connect timeout=None)')
问题核心是获取cl100k_base.tiktoken文件时网络超时,已尝试手动下载文件并设置TIKTOKEN_CACHE_DIR环境变量,但对需克隆使用的项目不够可靠,寻求无需下载文件到本地或仓库的解决方案。
解决方案
以下几种方案无需手动下载文件到本地或仓库,可解决该超时问题:
1. 显式指定分词器
在初始化AzureOpenAIEmbeddings时,显式指定对应分词器,避免自动触发外部文件下载:
import tiktoken from langchain_openai import AzureOpenAIEmbeddings # 显式获取text-embedding-ada-002对应的分词器 encoding = tiktoken.get_encoding("cl100k_base") # 初始化嵌入模型时传入分词器编码方法 self.model = AzureOpenAIEmbeddings( model=self.embedding_deployment, azure_endpoint=self.endpoint, openai_api_key=self.api_key, openai_api_version="2024-02-01", tokenizer=encoding.encode )
2. 升级依赖版本
部分旧版本的langchain或tiktoken在处理Azure部署模型时,存在分词器自动映射逻辑缺陷,升级到最新版本可修复该问题,避免触发外部下载请求:
pip install --upgrade langchain_openai langchain-community tiktoken
升级后直接初始化AzureOpenAIEmbeddings即可,无需额外配置分词器。
3. 配置网络代理(环境允许时)
如果项目运行环境有可用代理,为tiktoken设置代理后可正常获取外部文件:
import os # 替换为实际代理地址 os.environ["HTTP_PROXY"] = "http://your-proxy-address:port" os.environ["HTTPS_PROXY"] = "http://your-proxy-address:port" # 之后正常初始化AzureOpenAIEmbeddings和AzureSearch
4. 内部缓存分词器文件
将cl100k_base.tiktoken文件上传到内部存储服务,修改tiktoken加载逻辑从内部地址获取,无需用户手动下载:
import tiktoken from tiktoken.load import load_tiktoken_bpe from langchain_openai import AzureOpenAIEmbeddings # 替换为内部存储的文件地址 internal_encoding_path = "http://your-internal-storage/cl100k_base.tiktoken" # 手动构建分词器实例 encoding = tiktoken.Encoding( name="cl100k_base", pat_str=r"""'(?i:[sdmt]|ll|ve|re)|[^\r\n\p{L}\p{N}]?+\p{L}+|\p{N}{1,3}| ?[^\s\p{L}\p{N}]++[\r\n]*|\s*[\r\n]|\s+(?!\S)|\s+""", mergeable_ranks=load_tiktoken_bpe(internal_encoding_path), special_tokens={ "<|endoftext|>": 100257, "<|fim_prefix|>": 100258, "<|fim_middle|>": 100259, "<|fim_suffix|>": 100260, "<|endofprompt|>": 100276 } ) # 初始化嵌入模型时传入自定义分词器 self.model = AzureOpenAIEmbeddings( model=self.embedding_deployment, azure_endpoint=self.endpoint, openai_api_key=self.api_key, openai_api_version="2024-02-01", tokenizer=encoding.encode )
内容的提问来源于stack exchange,提问作者Evren Çetinkaya
相关产品推荐
相关产品推荐

