如何使用glob匹配两个文件夹中的同名文件并生成包含文件名与路径的DataFrame(Python/Bash实现)
Got it, here are two practical solutions to build that matching DataFrame—one using Python (with pandas) and another using Bash command-line tools. Both will pair files by their base filename (excluding extension) across your two folders.
Python Solution (using pandas)
This approach is flexible and easy to tweak for edge cases (like files with multiple dots in their names). First, install pandas if you haven’t already:
pip install pandas
Then use this script:
import os import pandas as pd # Replace these with your actual folder paths folder_original = "/home/dir/dir2" # Folder with original files (e.g., .docx) folder_modified = "/home/dir/dir1" # Folder with modified files (e.g., .csv) # Store files by their base name (without extension) original_files = {} modified_files = {} # Scan the original folder for root, _, files in os.walk(folder_original): for file in files: base_name = os.path.splitext(file)[0] original_files[base_name] = (file, root) # Scan the modified folder for root, _, files in os.walk(folder_modified): for file in files: base_name = os.path.splitext(file)[0] modified_files[base_name] = (file, root) # Only keep files that exist in both folders common_bases = set(original_files.keys()) & set(modified_files.keys()) # Build the DataFrame rows data = [] for base in common_bases: orig_file, orig_path = original_files[base] mod_file, mod_path = modified_files[base] data.append({ "file_name_original": orig_file, "file_path_original": orig_path, "file_name_modified": mod_file, "file_path_modified": mod_path }) # Create and display the final DataFrame df = pd.DataFrame(data) print(df) # Optional: Save to CSV for later use # df.to_csv("file_matches.csv", index=False)
Notes:
- If you want to include files that are missing in one folder (with
NaNvalues), replace thecommon_baseslogic withpd.mergeinstead. os.path.splitexthandles filenames with multiple dots (likereport.v2.docx) by taking everything before the last dot as the base name.
Bash Solution
If you prefer command-line tools, this script uses find, awk, sort, and join to generate the matching table without Python.
Save this as match_files.sh, make it executable (chmod +x match_files.sh), and run it:
#!/bin/bash # Replace these with your actual folder paths FOLDER_ORIGINAL="/home/dir/dir2" FOLDER_MODIFIED="/home/dir/dir1" # Generate temporary files with base name, filename, and directory path find "$FOLDER_ORIGINAL" -type f | awk -F/ '{ file = $NF; dir = substr($0, 1, length($0) - length(file) - 1); split(file, arr, "."); # Use this line for filenames with NO dots in the base (e.g., example.docx) base = arr[1]; # Uncomment below for filenames with dots in the base (e.g., report.v2.docx) # base = substr(file, 1, length(file) - length(arr[length(arr)]) - 1); print base "\t" file "\t" dir }' > original_temp.txt find "$FOLDER_MODIFIED" -type f | awk -F/ '{ file = $NF; dir = substr($0, 1, length($0) - length(file) - 1); split(file, arr, "."); base = arr[1]; # Uncomment below for dots in base names # base = substr(file, 1, length(file) - length(arr[length(arr)]) - 1); print base "\t" file "\t" dir }' > modified_temp.txt # Sort files by base name to prepare for joining sort original_temp.txt > original_sorted.txt sort modified_temp.txt > modified_sorted.txt # Join and format the output to match your desired columns echo -e "file_name_original\tfile_path_original\tfile_name_modified\tfile_path_modified" join -t $'\t' original_sorted.txt modified_sorted.txt | awk -F'\t' '{ print $2 "\t" $3 "\t" $5 "\t" $6 }' # Clean up temporary files rm original_temp.txt modified_temp.txt original_sorted.txt modified_sorted.txt
Notes:
- The output is tab-separated—you can convert it to CSV with
./match_files.sh | sed 's/\t/,/g' > file_matches.csv. - Switch to the alternative
baseline if your files have dots in their base names (likeproject.v1.pdf).
内容的提问来源于stack exchange,提问作者Noel Harris
相关产品推荐
相关产品推荐

