Ollama内部并行性验证及多线程代码优化咨询
问题解答
一、关于Ollama并行处理的理解纠正
你的理解不正确:
- Ollama处理单个请求时,主要依赖GPU完成模型运算,但CPU层面仅会用单个核心处理请求调度、数据预处理/后处理等辅助任务,所以htop看到单核心占用是正常的。
- 你用for循环串行发送请求时,Ollama会逐个处理任务,不会自动并行执行多个请求。这种情况下,GPU始终只在处理单个翻译任务,无法跑满H100的算力,因此V100和H100的执行时间差异不明显。
二、代码并行优化指导
针对你的多报告翻译场景(IO密集型请求),推荐两种并行实现方案:
方案1:多线程实现(基于ThreadPoolExecutor)
适合IO密集型任务,Python的GIL在网络等待时会释放,多线程能有效提升并发效率。
修改后的代码:
from functools import cached_property from concurrent.futures import ThreadPoolExecutor, as_completed from ollama import Client class TestOllama: @cached_property def ollama_client(self) -> Client: return Client(host="http://127.0.0.1:11434") def translate(self, text_to_translate: str): # 修正prompt方向:需求是英文转法文 ollama_response = self.ollama_client.generate( model="mistral", prompt=f"translate this English text into French: {text_to_translate}" ) return ollama_response['response'].lstrip(), ollama_response['total_duration'] def run(self): reports = ["reports_text_1", "reports_text_2"] # 替换为实际报告列表 # 控制并发数:根据GPU显存和Ollama承载能力调整,比如4-8 max_workers = 4 with ThreadPoolExecutor(max_workers=max_workers) as executor: # 提交所有翻译任务 future_to_report = {executor.submit(self.translate, report): report for report in reports} # 实时处理完成的任务 for future in as_completed(future_to_report): report = future_to_report[future] try: translated_report, total_duration = future.result() print(f"原报告摘要: {report[:50]}...\n翻译结果摘要: {translated_report[:50]}...\n耗时: {total_duration}\n") except Exception as e: print(f"处理报告时出错: {str(e)}") if __name__ == '__main__': job = TestOllama() job.run()
关键改动说明:
- 用
ThreadPoolExecutor创建线程池,通过max_workers控制并发请求数量,避免Ollama或GPU过载 - 用
as_completed实时获取已完成任务的结果,不用等待所有任务结束 - 修正原prompt的翻译方向错误(原代码写的是法译英,需求为英译法)
方案2:异步实现(基于AsyncClient)
Ollama官方提供异步客户端,适合高并发场景,效率比多线程更高。
修改后的代码:
import asyncio from functools import cached_property from ollama import AsyncClient class TestOllama: @cached_property def ollama_async_client(self) -> AsyncClient: return AsyncClient(host="http://127.0.0.1:11434") async def translate(self, text_to_translate: str): ollama_response = await self.ollama_async_client.generate( model="mistral", prompt=f"translate this English text into French: {text_to_translate}" ) return ollama_response['response'].lstrip(), ollama_response['total_duration'] async def run(self): reports = ["reports_text_1", "reports_text_2"] # 替换为实际报告列表 # 创建异步任务列表 tasks = [self.translate(report) for report in reports] # 并发执行所有任务,单个任务失败不影响其他任务 results = await asyncio.gather(*tasks, return_exceptions=True) # 批量处理结果 for idx, (result, report) in enumerate(zip(results, reports)): if isinstance(result, Exception): print(f"第{idx+1}份报告处理失败: {str(result)}") else: translated_report, total_duration = result print(f"第{idx+1}份报告\n原文本摘要: {report[:50]}...\n翻译结果摘要: {translated_report[:50]}...\n耗时: {total_duration}\n") if __name__ == '__main__': job = TestOllama() asyncio.run(job.run())
关键改动说明:
- 用
AsyncClient替代同步客户端,所有IO操作使用await关键字实现异步等待 - 用
asyncio.gather并发执行所有任务,return_exceptions=True保证单个任务失败不会中断整体流程 - 整体采用异步范式,适合高并发场景
三、额外优化建议
- 调整并发数:H100显存更大,可逐步测试调高并发数(比如8-16),找到既不触发显存溢出又能跑满GPU的最优值
- 模型适配:如果mistral对长文本支持有限,可尝试更大的模型(如llama3),但需注意显存占用
- 批量请求:若报告格式统一,可尝试将多份报告打包成单个请求(需调整prompt让模型批量翻译),减少请求开销,但要注意模型的上下文窗口限制
内容的提问来源于stack exchange,提问作者Januka samaranyake
相关产品推荐
相关产品推荐

