如何使用wget批量下载GitHub项目的全部原始格式文件?
Got it, let's fix this—your brace expansion trick works great for individual files, but when you need every raw file in a repo, here are two reliable methods to get the job done smoothly:
Method 1: GitHub API + Wget (Controlled, Preserves Directory Structure)
This approach lets you explicitly fetch all file paths first, then download each one while keeping the repo's folder hierarchy intact.
Fetch the full list of file paths
First, we'll use the GitHub API to pull all blob (file) paths from your target branch. You'll needjqinstalled (most systems have it available via package managers likeapt,brew, oryum) to parse the JSON response:curl -s https://api.github.com/repos/u/p/git/trees/b?recursive=1 | jq -r '.tree[] | select(.type=="blob") | .path' > file_list.txtReplace
uwith the repo owner's username,pwith the repo name, andbwith the branch name. This saves all valid file paths to a text file namedfile_list.txt.Batch download with Wget
Run this loop to download each file, automatically creating the same directory structure as the repo in your home folder:while read path; do wget -P ~/$(dirname "$path") "https://raw.githubusercontent.com/u/p/b/$path" done < file_list.txtThe
-Pflag tells Wget to create any missing directories needed to save each file in its correct location.
Note for large repos:
If you hit GitHub's API rate limit (common for big repos), add a personal access token to the curl command to increase your limit:
curl -s -H "Authorization: token YOUR_GITHUB_TOKEN" https://api.github.com/repos/u/p/git/trees/b?recursive=1 | jq -r '.tree[] | select(.type=="blob") | .path' > file_list.txt
You can generate a token in your GitHub account settings (no special scopes are needed for public repos).
Method 2: Recursive Wget (Quick & Simple)
If you don't want to mess with APIs, use Wget's recursive mode to crawl the raw file directory. This is perfect for small to medium repos:
wget -r -np -nH --cut-dirs=3 -P ~/ -A "*" https://raw.githubusercontent.com/u/p/b/
Let's break down the key flags:
-r: Enable recursive downloading-np: Don't climb to parent directories (avoids accidentally downloading other branches or repos)-nH: Skip creating a folder named after the GitHub domain--cut-dirs=3: Strip the first 3 directories (u/,p/,b/) so files save directly to your home folder with their original repo structure-P ~/: Set your home folder as the base save directory-A "*": Download all file types without restriction
Both methods will get you all raw files in their original format—pick whichever fits your workflow best!
内容的提问来源于stack exchange,提问作者user273993

