如何用Python直接获取GitHub目标文件夹,无需手动下载?
无需手动下载,直接从GitHub获取YAML数据的Python实现方法
核心思路
直接通过GitHub API遍历目标仓库的data目录结构,筛选出符合条件的YAML文件,然后请求文件的Raw内容进行解析,全程无需本地下载文件,就能获取实时更新的数据。
实现步骤与代码修改
需要先安装requests库(用于网络请求):
pip install requests
以下是修改后的完整代码,替换你原来的本地文件处理逻辑:
import yaml import requests from concurrent.futures import ThreadPoolExecutor # GitHub仓库相关配置 REPO_OWNER = "openstates" REPO_NAME = "people" TARGET_DIR = "data" def extractingOfficials(raw_url): try: # 发送请求获取Raw文件内容 response = requests.get(raw_url, timeout=10) response.raise_for_status() # 捕获HTTP请求错误 official = yaml.safe_load(response.text) return official except yaml.YAMLError as exc: print(f"解析YAML出错 {raw_url}: {exc}") except requests.exceptions.RequestException as exc: print(f"请求文件出错 {raw_url}: {exc}") return None def get_all_valid_yaml_urls(): valid_urls = [] # 递归遍历GitHub目录的函数 def traverse_dir(dir_path): api_url = f"https://api.github.com/repos/{REPO_OWNER}/{REPO_NAME}/contents/{dir_path}" try: response = requests.get(api_url, timeout=10) response.raise_for_status() items = response.json() for item in items: if item["type"] == "dir": # 跳过retired和committees目录 if "retired" in item["name"] or "committees" in item["name"]: continue # 递归遍历子目录 traverse_dir(item["path"]) elif item["type"] == "file": # 筛选符合条件的YAML文件 if item["name"].endswith(".yml") and "municipalities" not in item["name"]: valid_urls.append(item["download_url"]) except requests.exceptions.RequestException as exc: print(f"遍历目录出错 {dir_path}: {exc}") traverse_dir(TARGET_DIR) return valid_urls def mainExtractingFunction(): all_yml_urls = get_all_valid_yaml_urls() # 使用线程池并行处理请求 with ThreadPoolExecutor(max_workers=28) as pool: allofficials = list(pool.map(extractingOfficials, all_yml_urls)) # 过滤掉解析失败的None值 allofficials = [official for official in allofficials if official is not None] return allofficials # 调用示例 if __name__ == "__main__": officials = mainExtractingFunction() print(f"共获取到 {len(officials)} 条有效官员数据")
关键修改说明
- 替换本地文件读取为网络请求:
extractingOfficials函数不再接收本地文件路径,而是接收GitHub文件的Raw URL,通过requests.get获取内容后直接解析。 - 新增目录遍历逻辑:
get_all_valid_yaml_urls函数通过GitHub API递归遍历data目录,自动跳过retired、committees目录,筛选出符合要求的YAML文件的Raw地址。 - 增强异常处理:新增了网络请求的异常捕获,避免单个文件出错导致整个脚本崩溃。
- 保持并行处理效率:沿用原来的线程池方案,保证大量文件的处理效率。
备选方案(适合需要本地缓存的场景)
如果需要偶尔缓存数据避免频繁请求API,可以定时执行git clone或git pull命令拉取data目录,然后继续使用你原来的本地处理代码。示例命令:
# 首次克隆 git clone --depth 1 https://github.com/openstates/people.git temp_repo # 后续更新 cd temp_repo && git pull
之后将rootdir指向temp_repo/data即可。
内容的提问来源于stack exchange,提问作者Jimmy C
相关产品推荐
相关产品推荐

