Python使用asyncio并发转换xlsx到PDF时随机失败问题求助
问题背景
在Linux Mint系统下,使用Python调用soffice进行XLSX转PDF,顺序处理时全部正常,但通过asyncio并发处理(拆分文件列表为两半同时处理)后,100个文件中随机有约10个转换失败。
顺序处理代码
def convert_pdf_soffice(xlsx_file: str)->None: out_dir = './PdfDir/' print('Started conversion of ', xlsx_file) subprocess.run(['soffice', '--headless', '--convert-to', 'pdf', '--outdir', out_dir, xlsx_file]) print('Finished conversion of ', xlsx_file)
调用方式:
for file in xls_files_to_be_converted: convert_pdf_soffice(file)
并发处理代码
#!/usr/local/bin/python3 import os import asyncio import time async def convert_pdf_soffice(xlsx_files): out_dir = './PdfDir/' tasks = [] for xlsx_file in xlsx_files: print('Started conversion of ', xlsx_file) process = await asyncio.create_subprocess_exec( 'soffice', '--headless', '--convert-to', 'pdf', '--outdir', out_dir, xlsx_file, stdout=asyncio.subprocess.PIPE, stderr=asyncio.subprocess.PIPE ) tasks.append(process) for task, xlsx_file in zip(tasks, xlsx_files): stdout, stderr = await task.communicate() if task.returncode != 0: print(f'Conversion of {xlsx_file} failed with return code {task.returncode}') else: print('Finished conversion of ', xlsx_file) async def main(): start_t = time.time() INPUT_DIR = './XLSX/' OUTPUT_DIR = './PdfDir/' # Create folder if not exists if not os.path.exists(OUTPUT_DIR): os.makedirs(OUTPUT_DIR) # List of all xlsx files xlsx_file_list = [file for file in os.listdir(INPUT_DIR) if file.endswith('.xlsx')] # Split the list into two halves mid_index = len(xlsx_file_list) // 2 first_half = xlsx_file_list[:mid_index] second_half = xlsx_file_list[mid_index:] # List of xlsx files to be converted to pdf first_half_paths = [os.path.join(INPUT_DIR, file) for file in first_half] second_half_paths = [os.path.join(INPUT_DIR, file) for file in second_half] # Run conversions concurrently for both halves await asyncio.gather( convert_pdf_soffice(first_half_paths), convert_pdf_soffice(second_half_paths) ) end_t = time.time() duration_t = end_t - start_t print(f'Duration is {duration_t}') if __name__ == '__main__': asyncio.run(main())
故障原因分析
Soffice Headless模式的并发限制
LibreOffice/OpenOffice的--headless模式并非为高并发多实例设计,同时启动多个soffice进程时会触发资源竞争:比如进程尝试复用同一个UNO端口、临时目录文件锁冲突,或者共享的配置/缓存文件被多进程同时访问,导致部分进程启动失败或转换中断。无限制并发导致系统资源耗尽
当前并发代码会一次性启动所有文件对应的soffice进程,100个文件就会同时运行50+个soffice实例。每个soffice进程会占用大量内存(尤其是处理复杂Excel文件时),当系统内存、CPU或文件句柄耗尽时,操作系统会强制终止部分进程,导致转换失败。缺少错误细节排查
代码仅打印返回码,未输出stderr内容,无法得知具体失败原因(比如端口占用提示、内存不足报错、文件权限问题等),增加了定位难度。临时文件冲突
Soffice转换时会在临时目录生成中间文件,并发场景下可能出现不同进程生成同名临时文件,导致文件覆盖或读写权限冲突,进而转换失败。
解决方案
1. 限制并发数(快速修复)
使用asyncio.Semaphore控制同时运行的soffice实例数量,根据系统配置调整并发数(建议4-8个,避免资源过载):
#!/usr/local/bin/python3 import os import asyncio import time # 限制同时运行的soffice进程数 MAX_CONCURRENT = 4 semaphore = asyncio.Semaphore(MAX_CONCURRENT) async def convert_single_file(xlsx_file): out_dir = './PdfDir/' async with semaphore: print('Started conversion of ', xlsx_file) process = await asyncio.create_subprocess_exec( 'soffice', '--headless', '--convert-to', 'pdf', '--outdir', out_dir, xlsx_file, stdout=asyncio.subprocess.PIPE, stderr=asyncio.subprocess.PIPE ) stdout, stderr = await process.communicate() if process.returncode != 0: print(f'Conversion of {xlsx_file} failed with return code {process.returncode}') # 打印错误详情 print(f'Stderr: {stderr.decode("utf-8")}') else: print('Finished conversion of ', xlsx_file) async def main(): start_t = time.time() INPUT_DIR = './XLSX/' OUTPUT_DIR = './PdfDir/' if not os.path.exists(OUTPUT_DIR): os.makedirs(OUTPUT_DIR) xlsx_file_list = [os.path.join(INPUT_DIR, file) for file in os.listdir(INPUT_DIR) if file.endswith('.xlsx')] # 创建所有转换任务 tasks = [convert_single_file(file) for file in xlsx_file_list] await asyncio.gather(*tasks) end_t = time.time() duration_t = end_t - start_t print(f'Duration is {duration_t}') if __name__ == '__main__': asyncio.run(main())
2. 复用Soffice实例(高效最优解)
启动一个常驻的soffice进程,通过UNO接口发送转换请求,避免频繁启动进程的开销和资源冲突:
# 示例:使用pyoo库连接常驻soffice进程 # 先安装pyoo: pip install pyoo import os import pyoo # 先在终端启动常驻soffice进程 # soffice --headless --accept="socket,host=localhost,port=2002;urp;" --norestore def convert_with_uno(xlsx_file, out_dir): desktop = pyoo.Desktop('localhost', 2002) doc = desktop.open(xlsx_file) try: doc.export_to_pdf(os.path.join(out_dir, os.path.basename(xlsx_file).replace('.xlsx', '.pdf'))) finally: doc.close() # 调用时可结合线程池/asyncio进行并发(注意UNO线程安全)
3. 增加错误日志
在失败时打印stderr内容,帮助定位具体问题,比如端口占用、文件损坏等。
4. 系统资源优化
- 确保系统有足够内存(可通过
free -h查看),避免因内存不足导致进程被OOM killer终止。 - 调整系统文件句柄限制(
ulimit -n),避免因打开文件过多报错。
内容的提问来源于stack exchange,提问作者Aden Denitz

