使用subprocess.PIPE致Python subprocess运行缓慢,求优化方案
问题:subprocess.Popen 重定向stdout到PIPE导致性能骤降的解决办法
我通过多进程结合ssh批量获取多台主机的指定目录,发现使用subprocess.Popen并设置stdout=subprocess.PIPE时,运行时间显著增加。希望找到既能获取命令输出,又能汇总存在匹配目录主机的高效方法。
性能极差的代码(带stdout=PIPE)
processes = [] max_hosts = 15 for host in hosts: cmd = """ssh {} -o LogLevel=Error 'ls -ld /home/*/ 2>/dev/null | egrep SOMETEXT'""".format(host) p = subprocess.Popen(cmd, stdout=subprocess.PIPE, shell=True) print(p.stdout.read()) processes.add(p) if len(processes) >= max_hosts: os.wait() processes.difference_update([p for p in processes if p.poll() is not None]) for p in processes: if p.poll() is None: p.wait()
耗时:
real 0m22.965s user 0m2.456s sys 0m5.701s
性能正常的代码(无stdout重定向)
processes = [] max_hosts = 15 for host in hosts: cmd = """ssh {} -o LogLevel=Error 'ls -ld /home/*/ 2>/dev/null | egrep SOMETEXT'""".format(host) p = subprocess.Popen(cmd, shell=True) processes.add(p) if len(processes) >= max_hosts: os.wait() processes.difference_update([p for p in processes if p.poll() is not None]) for p in processes: if p.poll() is None: p.wait()
耗时:
real 0m1.337s user 0m2.621s sys 0m6.186s
解决办法
1. 避免启动进程后立刻阻塞读取stdout
你当前代码性能暴跌的核心原因不是stdout=PIPE本身,而是启动进程后马上调用p.stdout.read()阻塞了主线程,把并发执行变成了串行执行。正确的做法是先批量启动并发进程,之后再统一读取输出:
import subprocess import os processes = [] max_hosts = 15 host_map = {} # 关联进程与对应主机 for host in hosts: cmd = """ssh {} -o LogLevel=Error 'ls -ld /home/*/ 2>/dev/null | egrep SOMETEXT'""".format(host) p = subprocess.Popen(cmd, stdout=subprocess.PIPE, shell=True) host_map[p] = host processes.append(p) # 控制并发数,处理已完成的进程 if len(processes) >= max_hosts: for p in processes[:]: if p.poll() is not None: output = p.stdout.read().decode('utf-8') print(f"[{host_map[p]}]\n{output}") processes.remove(p) del host_map[p] # 等待任意子进程结束 if processes: os.wait() # 处理剩余未完成的进程 for p in processes: p.wait() output = p.stdout.read().decode('utf-8') print(f"[{host_map[p]}]\n{output}")
2. 用并发库简化逻辑(推荐)
使用concurrent.futures.ThreadPoolExecutor可以更简洁地管理并发,同时通过capture_output=True获取输出,性能不受影响:
from concurrent.futures import ThreadPoolExecutor import subprocess def check_host(host): cmd = """ssh {} -o LogLevel=Error 'ls -ld /home/*/ 2>/dev/null | egrep SOMETEXT'""".format(host) result = subprocess.run(cmd, shell=True, capture_output=True, text=True) return (host, result.stdout.strip()) if result.stdout else (host, None) # 控制并发数为15 with ThreadPoolExecutor(max_workers=15) as executor: results = executor.map(check_host, hosts) # 汇总并打印结果 for host, output in results: if output: print(f"主机 {host} 存在匹配目录:\n{output}")
3. 优化远程命令减少数据传输
如果核心需求是汇总存在匹配目录的主机,可以让远程命令只返回状态(减少数据传输),再按需获取完整输出:
# 远程命令用egrep -q判断是否存在匹配,不输出内容 def check_host_exist(host): cmd = """ssh {} -o LogLevel=Error 'ls -ld /home/*/ 2>/dev/null | egrep -q SOMETEXT'""".format(host) return host, subprocess.run(cmd, shell=True).returncode == 0 # 先批量检查哪些主机有匹配 with ThreadPoolExecutor(max_workers=15) as executor: exist_results = executor.map(check_host_exist, hosts) # 只给有匹配的主机获取完整输出 target_hosts = [host for host, exist in exist_results if exist] with ThreadPoolExecutor(max_workers=15) as executor: output_results = executor.map(check_host, target_hosts) # 打印最终结果 for host, output in output_results: print(f"主机 {host} 匹配目录详情:\n{output}")
核心原因总结
stdout=subprocess.PIPE本身不会导致性能下降,问题出在启动进程后立刻阻塞读取输出,破坏了多进程并发的逻辑。只要让进程真正并行执行,再统一处理输出,就能恢复正常性能。
内容的提问来源于stack exchange,提问作者user14069885
相关产品推荐
相关产品推荐

