如何在CodeCommit推送事件触发时自动上传全量文件至S3?
解决CodeCommit全量文件推送至S3的问题
我来帮你搞定这个问题!你的代码目前有两个核心问题导致了当前的困境:一是依赖get_differences只能获取变更文件,无法拿到仓库全量内容;二是S3上传时没有传入压缩包的内容,直接导致了0字节的文件。下面我会一步步帮你修正,实现推送事件触发时同步全量文件到S3的功能。
问题分析
- 仅获取变更文件:
get_differencesAPI只能返回两次提交之间的差异文件,当没有新变更时,这个列表是空的,自然生成的tar包也没有内容。 - S3上传遗漏内容:你当前的
put_object调用没有传入Body参数,相当于只在S3创建了一个空文件,这就是压缩包显示0字节的直接原因。
解决方案:获取仓库全量文件
要获取CodeCommit仓库的所有文件,我们需要递归遍历仓库的文件夹结构,使用get_folder_contents API来获取每个文件夹下的文件和子文件夹,然后逐一获取每个文件的Blob内容。
修改后的完整代码
import boto3 import pathlib import tarfile import io import sys import os codecommit = boto3.client("codecommit") def get_all_files(repository_name, branch="master", folder_path=""): """递归获取仓库指定分支下的所有文件""" all_files = [] response = codecommit.get_folder_contents( repositoryName=repository_name, commitSpecifier=branch, folderPath=folder_path ) # 处理当前文件夹下的文件 for item in response.get("files", []): all_files.append({ "path": item["filePath"], "blobId": item["blobId"], "mode": item["fileMode"] }) # 递归处理子文件夹 for folder in response.get("folders", []): all_files.extend(get_all_files( repository_name, branch, folder["folderPath"] )) # 处理分页(如果文件数量多的话) while "nextToken" in response: response = codecommit.get_folder_contents( repositoryName=repository_name, commitSpecifier=branch, folderPath=folder_path, nextToken=response["nextToken"] ) # 添加当前页的文件 for item in response.get("files", []): all_files.append({ "path": item["filePath"], "blobId": item["blobId"], "mode": item["fileMode"] }) # 递归处理子文件夹 for folder in response.get("folders", []): all_files.extend(get_all_files( repository_name, branch, folder["folderPath"] )) return all_files def main(event, context): repository_name = event['Records'][0]['eventSourceARN'].split(':')[5] repository_path = pathlib.Path(repository_name) buf = io.BytesIO() with tarfile.open(None, "w:gz", fileobj=buf) as tar: # 获取仓库全量文件 all_files = get_all_files(repository_name) for file_info in all_files: blobid = file_info["blobId"] path = file_info["path"] # 获取文件Blob内容 blob = codecommit.get_blob( repositoryName=repository_name, blobId=blobid ) # 构建TarInfo对象,设置文件路径和大小 tarinfo = tarfile.TarInfo(str(repository_path / path)) tarinfo.size = len(blob["content"]) # 添加文件到压缩包 tar.addfile(tarinfo, io.BytesIO(blob["content"])) # 获取压缩包的字节内容 tarobject = buf.getvalue() # 上传到S3 s3 = boto3.client('s3') s3_bucket = 'csharp-codes-codecommit' # S3键建议用仓库名+压缩包后缀,不要用/tmp/(S3没有本地路径概念) s3_key = f"{repository_name}.tar.gz" response = s3.put_object( Bucket=s3_bucket, Key=s3_key, Body=tarobject # 关键:传入压缩包内容 ) print(f"Successfully uploaded {len(tarobject)} bytes to S3: {s3_key}") return response
关键修改点说明
- 全量文件获取:新增
get_all_files函数,递归遍历仓库的所有文件夹和文件,确保拿到完整的仓库内容。 - 修复S3上传:在
put_object中添加Body=tarobject参数,把生成的压缩包内容上传到S3。同时修正了S3的键名,避免使用本地路径格式(S3的键是逻辑路径,不需要/tmp/)。 - 鲁棒性处理:添加了分页处理,即使仓库文件数量超过API单次返回限制,也能完整获取所有文件。
这样修改后,每次推送事件触发时,Lambda都会拉取指定分支的全量文件,打包后上传到S3,再也不会出现0字节的压缩包了。
内容的提问来源于stack exchange,提问作者sarojini hiregoudar
相关产品推荐
相关产品推荐

