subprocess.Popen/run未等待aws s3 sync命令执行完成问题
解决Python调用aws s3 sync未完成下载就执行后续代码的问题
你的核心问题是:调用AWS CLI的s3 sync命令时,进程因stdout/stderr缓冲区填满而阻塞,表现为仅下载少量文件后代码就继续执行,而非真正完成同步。
问题根源
当使用subprocess.PIPE捕获命令输出时,AWS CLI同步大量文件会产生大量输出,填满PIPE的缓冲区后,进程会被阻塞无法继续执行下载操作。你原代码中的while process.poll() is None空循环不仅多余,还会让进程一直卡在缓冲阻塞状态,导致你误以为同步提前结束。
修正方案
方案1:不捕获输出,避免缓冲阻塞
直接让AWS CLI的输出继承父进程的终端流,不会产生缓冲阻塞问题,同时subprocess.run()默认会等待进程完全结束:
def sync_s3(self, aws_cli_path, s3_path, local_directory, aws_profile): print("Running sync command from s3") try: result = subprocess.run( [aws_cli_path, "s3", "sync", s3_path, local_directory, "--profile", aws_profile], check=True ) print("S3 sync command completed successfully") logging.info('S3 sync completed successfully') except subprocess.CalledProcessError as e: error_msg = e.stderr.decode() if e.stderr else 'Unknown error' print(f"Command failed with error: {error_msg}") logging.critical(f'S3 sync failed: {error_msg}') sys.exit(701)
方案2:捕获输出但正确处理缓冲
如果需要记录命令输出,使用communicate()方法,它会自动等待进程结束并读取所有输出,避免缓冲阻塞:
def sync_s3(self, aws_cli_path, s3_path, local_directory, aws_profile): print("Running sync command from s3") try: process = subprocess.Popen( [aws_cli_path, "s3", "sync", s3_path, local_directory, "--profile", aws_profile], stdout=subprocess.PIPE, stderr=subprocess.PIPE, text=True ) # communicate()会等待进程终止并读取全部输出 stdout, stderr = process.communicate() if process.returncode != 0: raise Exception(stderr) else: print(f"S3 sync command completed successfully: {stdout}") logging.info(f'S3 sync completed successfully: {stdout}') except Exception as e: print(f"Command failed with error: {e}") logging.critical(f'S3 sync failed: {e}') sys.exit(701)
补充说明
如果你之前使用subprocess.run()也出现相同问题,说明当时你同样设置了stdout=PIPE但未处理输出,导致缓冲阻塞。改用上述两种方案即可解决。
内容的提问来源于stack exchange,提问作者suaveplant
相关产品推荐
相关产品推荐

