You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python直接获取GitHub目标文件夹,无需手动下载?

无需手动下载,直接从GitHub获取YAML数据的Python实现方法

核心思路

直接通过GitHub API遍历目标仓库的data目录结构,筛选出符合条件的YAML文件,然后请求文件的Raw内容进行解析,全程无需本地下载文件,就能获取实时更新的数据。

实现步骤与代码修改

需要先安装requests库(用于网络请求):

pip install requests

以下是修改后的完整代码,替换你原来的本地文件处理逻辑:

import yaml
import requests
from concurrent.futures import ThreadPoolExecutor

# GitHub仓库相关配置
REPO_OWNER = "openstates"
REPO_NAME = "people"
TARGET_DIR = "data"

def extractingOfficials(raw_url):
    try:
        # 发送请求获取Raw文件内容
        response = requests.get(raw_url, timeout=10)
        response.raise_for_status()  # 捕获HTTP请求错误
        official = yaml.safe_load(response.text)
        return official
    except yaml.YAMLError as exc:
        print(f"解析YAML出错 {raw_url}: {exc}")
    except requests.exceptions.RequestException as exc:
        print(f"请求文件出错 {raw_url}: {exc}")
    return None

def get_all_valid_yaml_urls():
    valid_urls = []
    # 递归遍历GitHub目录的函数
    def traverse_dir(dir_path):
        api_url = f"https://api.github.com/repos/{REPO_OWNER}/{REPO_NAME}/contents/{dir_path}"
        try:
            response = requests.get(api_url, timeout=10)
            response.raise_for_status()
            items = response.json()
            for item in items:
                if item["type"] == "dir":
                    # 跳过retired和committees目录
                    if "retired" in item["name"] or "committees" in item["name"]:
                        continue
                    # 递归遍历子目录
                    traverse_dir(item["path"])
                elif item["type"] == "file":
                    # 筛选符合条件的YAML文件
                    if item["name"].endswith(".yml") and "municipalities" not in item["name"]:
                        valid_urls.append(item["download_url"])
        except requests.exceptions.RequestException as exc:
            print(f"遍历目录出错 {dir_path}: {exc}")
    
    traverse_dir(TARGET_DIR)
    return valid_urls

def mainExtractingFunction():
    all_yml_urls = get_all_valid_yaml_urls()
    # 使用线程池并行处理请求
    with ThreadPoolExecutor(max_workers=28) as pool:
        allofficials = list(pool.map(extractingOfficials, all_yml_urls))
        # 过滤掉解析失败的None值
        allofficials = [official for official in allofficials if official is not None]
    return allofficials

# 调用示例
if __name__ == "__main__":
    officials = mainExtractingFunction()
    print(f"共获取到 {len(officials)} 条有效官员数据")

关键修改说明

  1. 替换本地文件读取为网络请求:extractingOfficials函数不再接收本地文件路径,而是接收GitHub文件的Raw URL,通过requests.get获取内容后直接解析。
  2. 新增目录遍历逻辑:get_all_valid_yaml_urls函数通过GitHub API递归遍历data目录,自动跳过retired、committees目录,筛选出符合要求的YAML文件的Raw地址。
  3. 增强异常处理:新增了网络请求的异常捕获,避免单个文件出错导致整个脚本崩溃。
  4. 保持并行处理效率:沿用原来的线程池方案,保证大量文件的处理效率。

备选方案(适合需要本地缓存的场景)

如果需要偶尔缓存数据避免频繁请求API,可以定时执行git clone或git pull命令拉取data目录,然后继续使用你原来的本地处理代码。示例命令:

# 首次克隆
git clone --depth 1 https://github.com/openstates/people.git temp_repo
# 后续更新
cd temp_repo && git pull

之后将rootdir指向temp_repo/data即可。

内容的提问来源于stack exchange,提问作者Jimmy C

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.13 20:35:03