如何用Python从Git仓库获取同一CSV文件的两个不同版本?
用Python从Git仓库获取CSV文件的指定版本
当然可以实现,下面提供两种常用方案,直接获取指定版本的CSV内容并处理,无需先保存为本地文件:
方法1:调用Git命令行(无需额外依赖)
利用Git自带的git show命令提取指定版本的文件内容,通过Python的subprocess模块执行命令,再用io.StringIO将内容转换成可读取的文件对象,直接用于csv.DictReader处理:
import subprocess import csv from io import StringIO import sys # 仓库根目录和目标文件路径 repo_root = "C:\\Users\\Desktop" csv_file_path = "the_csv_file.csv" # 目标版本的哈希值 version_3_hash = "26ff7u90fa988975b8e94bb7ca1ecee94ed8ea64" version_2_hash = "b1c9b1200382b012c4c30284f2322bbb9c81a15f" def get_git_file_content(commit_hash, file_path, repo_dir): # 执行git show命令获取指定版本的文件内容 result = subprocess.run( ["git", "-C", repo_dir, "show", f"{commit_hash}:{file_path}"], capture_output=True, text=True, encoding="utf-8" # 根据你的CSV文件编码调整 ) if result.returncode != 0: print(f"获取文件失败: {result.stderr}", file=sys.stderr) return None # 将字符串内容转为文件对象 return StringIO(result.stdout) # 获取两个版本的文件对象 file_v3 = get_git_file_content(version_3_hash, csv_file_path, repo_root) file_v2 = get_git_file_content(version_2_hash, csv_file_path, repo_root) if file_v3 and file_v2: # 用csv.DictReader解析内容 csv_v3 = csv.DictReader(file_v3) csv_v2 = csv.DictReader(file_v2) # 示例:遍历版本3的所有行 for row in csv_v3: print(row) # 显式关闭文件对象(可选,StringIO会自动回收) file_v3.close() file_v2.close()
方法2:使用GitPython库(更Pythonic)
如果需要频繁和Git仓库交互,推荐使用GitPython库,它封装了Git操作,代码更简洁易读:
第一步:安装依赖
pip install gitpython
第二步:实现代码
import git import csv from io import StringIO # 打开目标Git仓库 repo = git.Repo("C:\\Users\\Desktop") # 目标版本哈希和文件路径 version_3_hash = "26ff7u90fa988975b8e94bb7ca1ecee94ed8ea64" version_2_hash = "b1c9b1200382b012c4c30284f2322bbb9c81a15f" csv_file_path = "the_csv_file.csv" def get_versioned_file(repo, commit_hash, file_path): # 获取指定提交对象 commit = repo.commit(commit_hash) # 从提交的树结构中找到目标文件 file_blob = commit.tree / file_path # 读取文件内容并转为文件对象 return StringIO(file_blob.data_stream.read().decode("utf-8")) # 调整编码 # 获取两个版本的文件对象 file_v3 = get_versioned_file(repo, version_3_hash, csv_file_path) file_v2 = get_versioned_file(repo, version_2_hash, csv_file_path) if file_v3 and file_v2: # 解析CSV内容 csv_v3 = csv.DictReader(file_v3) csv_v2 = csv.DictReader(file_v2) # 示例:遍历版本2的所有行 for row in csv_v2: print(row) file_v3.close() file_v2.close()
两种方案都可以实现你的需求,无需将文件保存到本地再读取,直接在内存中完成内容获取和解析。
内容的提问来源于stack exchange,提问作者user3084383
相关产品推荐
相关产品推荐

