Git高效获取文件最后提交信息:含短消息、日期及作者
Great question—dealing with large repos (16k files!) means we need to avoid per-file Git commands, which add up to thousands of slow process spawns. Let's break down why your initial approaches struggled, then fix it with a single, fast Git command paired with a simple processing script.
Why Your Initial Methods Were Slow
- The loop over each file runs
git logonce per file. For 16k files, that's 16k separate Git processes—each one has to initialize and traverse the commit graph, which is why it's so slow. - The fast
git log -- */*command gives you commit metadata but doesn't map each file to its latest commit, so you can't directly get the per-file info you need.
The Efficient Solution
We'll use a single git log command to output all commit details along with the files each commit touched, then process that output to track the latest commit for every file. This way, Git only traverses the commit graph once, which is optimized and fast even for huge repos.
Step 1: Run the Optimized Git Command
This command outputs commit markers with metadata, followed by the files changed in each commit:
git log --pretty=format:"COMMIT|||%h|||%an|||%ai|||%s" --name-only
Let's break down the format placeholders:
%h: Short commit hash (7 characters, like GitHub uses)%an: Author name%ai: ISO 8601 date (with timezone, we'll trim that later if needed)%s: Commit subject (short message)COMMIT|||: A unique marker to separate commit entries from file paths
Step 2: Process the Output with a Bash Script
Save this as get_last_commits.sh (make it executable with chmod +x get_last_commits.sh):
#!/bin/bash # Associative array to track the latest commit for each file declare -A file_commits # Get total number of files to enable early exit (optional but helpful) total_files=$(git ls-files | wc -l) # Process git log output line by line git log --pretty=format:"COMMIT|||%h|||%an|||%ai|||%s" --name-only | while IFS= read -r line; do if [[ $line == COMMIT|||* ]]; then # Split commit info using our unique delimiter IFS='|||' read -r _ commit_hash author date subject <<< "$line" # Format date to remove timezone (matches GitHub's display) formatted_date=${date% *} current_commit_details="Commit: $commit_hash\nAuthor: $author\nDate: $formatted_date\nMessage: $subject" else # Skip empty lines (Git inserts these between commits and file lists) [[ -z $line ]] && continue # Only record the first (newest) commit we see for each file if [[ ! ${file_commits[$line]} ]]; then file_commits[$line]="$current_commit_details" # Optional: Print progress as we process files echo -e "$line\n$current_commit_details\n---" # Exit early once all files have their latest commit if [[ ${#file_commits[@]} -eq $total_files ]]; then break fi fi fi done # Optional: Print all results sorted by file path echo -e "\n=== All Files with Latest Commit Details (Sorted) ===" for file in $(printf "%s\n" "${!file_commits[@]}" | sort); do echo -e "$file\n${file_commits[$file]}\n---" done
How This Works
- Single Git Process: We run
git logonce, which efficiently traverses the commit graph from newest to oldest. - Track Latest Commit: For each commit, we check all files it changed. If a file hasn't been recorded yet, we save this commit as its latest one (since we're processing newest first).
- Early Exit: Once we've recorded the latest commit for every file, we stop processing older commits—saving even more time.
Performance
For a 16k file repo, this should take just a few seconds (far faster than your initial 20 seconds for 500 files). The exact time depends on your repo's commit history size, but it's orders of magnitude better than per-file commands.
Notes for Git 2.14
All the flags used here (--pretty=format, --name-only, %h, %an, %ai, %s) are fully supported in Git 2.14, so you won't run into compatibility issues.
内容的提问来源于stack exchange,提问作者Mike

