You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从Git仓库获取同一CSV文件的两个不同版本?

用Python从Git仓库获取CSV文件的指定版本

当然可以实现,下面提供两种常用方案,直接获取指定版本的CSV内容并处理,无需先保存为本地文件:

方法1:调用Git命令行(无需额外依赖)

利用Git自带的git show命令提取指定版本的文件内容,通过Python的subprocess模块执行命令,再用io.StringIO将内容转换成可读取的文件对象,直接用于csv.DictReader处理:

import subprocess
import csv
from io import StringIO
import sys

# 仓库根目录和目标文件路径
repo_root = "C:\\Users\\Desktop"
csv_file_path = "the_csv_file.csv"

# 目标版本的哈希值
version_3_hash = "26ff7u90fa988975b8e94bb7ca1ecee94ed8ea64"
version_2_hash = "b1c9b1200382b012c4c30284f2322bbb9c81a15f"

def get_git_file_content(commit_hash, file_path, repo_dir):
    # 执行git show命令获取指定版本的文件内容
    result = subprocess.run(
        ["git", "-C", repo_dir, "show", f"{commit_hash}:{file_path}"],
        capture_output=True,
        text=True,
        encoding="utf-8"  # 根据你的CSV文件编码调整
    )
    if result.returncode != 0:
        print(f"获取文件失败: {result.stderr}", file=sys.stderr)
        return None
    # 将字符串内容转为文件对象
    return StringIO(result.stdout)

# 获取两个版本的文件对象
file_v3 = get_git_file_content(version_3_hash, csv_file_path, repo_root)
file_v2 = get_git_file_content(version_2_hash, csv_file_path, repo_root)

if file_v3 and file_v2:
    # 用csv.DictReader解析内容
    csv_v3 = csv.DictReader(file_v3)
    csv_v2 = csv.DictReader(file_v2)
    
    # 示例:遍历版本3的所有行
    for row in csv_v3:
        print(row)
    
    # 显式关闭文件对象(可选,StringIO会自动回收)
    file_v3.close()
    file_v2.close()

方法2:使用GitPython库(更Pythonic)

如果需要频繁和Git仓库交互,推荐使用GitPython库,它封装了Git操作,代码更简洁易读:

第一步:安装依赖

pip install gitpython

第二步:实现代码

import git
import csv
from io import StringIO

# 打开目标Git仓库
repo = git.Repo("C:\\Users\\Desktop")

# 目标版本哈希和文件路径
version_3_hash = "26ff7u90fa988975b8e94bb7ca1ecee94ed8ea64"
version_2_hash = "b1c9b1200382b012c4c30284f2322bbb9c81a15f"
csv_file_path = "the_csv_file.csv"

def get_versioned_file(repo, commit_hash, file_path):
    # 获取指定提交对象
    commit = repo.commit(commit_hash)
    # 从提交的树结构中找到目标文件
    file_blob = commit.tree / file_path
    # 读取文件内容并转为文件对象
    return StringIO(file_blob.data_stream.read().decode("utf-8"))  # 调整编码

# 获取两个版本的文件对象
file_v3 = get_versioned_file(repo, version_3_hash, csv_file_path)
file_v2 = get_versioned_file(repo, version_2_hash, csv_file_path)

if file_v3 and file_v2:
    # 解析CSV内容
    csv_v3 = csv.DictReader(file_v3)
    csv_v2 = csv.DictReader(file_v2)
    
    # 示例:遍历版本2的所有行
    for row in csv_v2:
        print(row)
    
    file_v3.close()
    file_v2.close()

两种方案都可以实现你的需求,无需将文件保存到本地再读取,直接在内存中完成内容获取和解析。

内容的提问来源于stack exchange,提问作者user3084383

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 19:17:06