Python3中curl执行滞后导致后续指令失效的解决方法咨询
解决Python3中命令异步执行导致的文本提取失效问题
问题根源
你遇到的问题是因为os.popen在Python3中是异步启动子进程的,代码会直接往下执行而不等待curl下载完成。Python2中因解释器行为差异,刚好能保证curl执行完再跑grep,但Python3里这个逻辑不再成立,导致grep在temptext.txt还没写完时就执行,最终output.txt为空。
解决方案
方案1:用subprocess.run(推荐,Python3官方替代方案)
subprocess.run默认会阻塞等待子进程执行完毕,完美保证命令执行顺序:
#!/usr/bin/python3 import subprocess # 等待curl下载完成,将结果写入temptext.txt subprocess.run( ['curl', 'https://www.york.ac.uk/teaching/cws/wws/webpage1.html', '-o', 'temptext.txt'], check=True # 如果curl执行失败会抛出异常,便于排查问题 ) # 执行grep并将结果写入output.txt subprocess.run( ['grep', 'understanding', 'temptext.txt'], stdout=open('output.txt', 'w') )
方案2:优化版——不用临时文件,直接管道传输
省去临时文件的IO操作,效率更高:
#!/usr/bin/python3 import subprocess with open('output.txt', 'w') as output_file: # 启动curl进程,将输出通过管道传给grep curl_process = subprocess.Popen( ['curl', 'https://www.york.ac.uk/teaching/cws/wws/webpage1.html'], stdout=subprocess.PIPE ) # 启动grep进程,读取curl的输出并写入文件 subprocess.run( ['grep', 'understanding'], stdin=curl_process.stdout, stdout=output_file ) # 等待curl进程结束 curl_process.wait()
方案3:兼容旧代码——强制os.popen等待进程结束
如果一定要保留os.popen的写法,需要通过读取输出或关闭文件对象来等待进程完成:
#!/usr/bin/python3 import os # 执行curl并等待完成 curl_proc = os.popen('curl https://www.york.ac.uk/teaching/cws/wws/webpage1.html > temptext.txt') curl_proc.read() # 读取输出会阻塞到进程结束 curl_proc.close() # 执行grep并等待完成 grep_proc = os.popen('grep understanding temptext.txt | head -n 1 > output.txt') grep_proc.read() grep_proc.close()
方案4:用shell管道直接串联命令
如果命令是固定的,也可以直接让shell处理管道顺序:
#!/usr/bin/python3 import subprocess subprocess.run( 'curl https://www.york.ac.uk/teaching/cws/wws/webpage1.html | grep understanding | head -n 1 > output.txt', shell=True, check=True )
注意:
shell=True存在安全风险,如果命令中包含用户输入内容不建议使用,但你的场景中命令固定,所以可以放心用。
内容的提问来源于stack exchange,提问作者Daz Voz
相关产品推荐
相关产品推荐

