You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用glob匹配两个文件夹中的同名文件并生成包含文件名与路径的DataFrame(Python/Bash实现)

Got it, here are two practical solutions to build that matching DataFrame—one using Python (with pandas) and another using Bash command-line tools. Both will pair files by their base filename (excluding extension) across your two folders.


Python Solution (using pandas)

This approach is flexible and easy to tweak for edge cases (like files with multiple dots in their names). First, install pandas if you haven’t already:

pip install pandas

Then use this script:

import os
import pandas as pd

# Replace these with your actual folder paths
folder_original = "/home/dir/dir2"  # Folder with original files (e.g., .docx)
folder_modified = "/home/dir/dir1"  # Folder with modified files (e.g., .csv)

# Store files by their base name (without extension)
original_files = {}
modified_files = {}

# Scan the original folder
for root, _, files in os.walk(folder_original):
    for file in files:
        base_name = os.path.splitext(file)[0]
        original_files[base_name] = (file, root)

# Scan the modified folder
for root, _, files in os.walk(folder_modified):
    for file in files:
        base_name = os.path.splitext(file)[0]
        modified_files[base_name] = (file, root)

# Only keep files that exist in both folders
common_bases = set(original_files.keys()) & set(modified_files.keys())

# Build the DataFrame rows
data = []
for base in common_bases:
    orig_file, orig_path = original_files[base]
    mod_file, mod_path = modified_files[base]
    data.append({
        "file_name_original": orig_file,
        "file_path_original": orig_path,
        "file_name_modified": mod_file,
        "file_path_modified": mod_path
    })

# Create and display the final DataFrame
df = pd.DataFrame(data)
print(df)

# Optional: Save to CSV for later use
# df.to_csv("file_matches.csv", index=False)

Notes:

  • If you want to include files that are missing in one folder (with NaN values), replace the common_bases logic with pd.merge instead.
  • os.path.splitext handles filenames with multiple dots (like report.v2.docx) by taking everything before the last dot as the base name.

Bash Solution

If you prefer command-line tools, this script uses find, awk, sort, and join to generate the matching table without Python.

Save this as match_files.sh, make it executable (chmod +x match_files.sh), and run it:

#!/bin/bash

# Replace these with your actual folder paths
FOLDER_ORIGINAL="/home/dir/dir2"
FOLDER_MODIFIED="/home/dir/dir1"

# Generate temporary files with base name, filename, and directory path
find "$FOLDER_ORIGINAL" -type f | awk -F/ '{
    file = $NF;
    dir = substr($0, 1, length($0) - length(file) - 1);
    split(file, arr, ".");
    # Use this line for filenames with NO dots in the base (e.g., example.docx)
    base = arr[1];
    # Uncomment below for filenames with dots in the base (e.g., report.v2.docx)
    # base = substr(file, 1, length(file) - length(arr[length(arr)]) - 1);
    print base "\t" file "\t" dir
}' > original_temp.txt

find "$FOLDER_MODIFIED" -type f | awk -F/ '{
    file = $NF;
    dir = substr($0, 1, length($0) - length(file) - 1);
    split(file, arr, ".");
    base = arr[1];
    # Uncomment below for dots in base names
    # base = substr(file, 1, length(file) - length(arr[length(arr)]) - 1);
    print base "\t" file "\t" dir
}' > modified_temp.txt

# Sort files by base name to prepare for joining
sort original_temp.txt > original_sorted.txt
sort modified_temp.txt > modified_sorted.txt

# Join and format the output to match your desired columns
echo -e "file_name_original\tfile_path_original\tfile_name_modified\tfile_path_modified"
join -t $'\t' original_sorted.txt modified_sorted.txt | awk -F'\t' '{
    print $2 "\t" $3 "\t" $5 "\t" $6
}'

# Clean up temporary files
rm original_temp.txt modified_temp.txt original_sorted.txt modified_sorted.txt

Notes:

  • The output is tab-separated—you can convert it to CSV with ./match_files.sh | sed 's/\t/,/g' > file_matches.csv.
  • Switch to the alternative base line if your files have dots in their base names (like project.v1.pdf).

内容的提问来源于stack exchange,提问作者Noel Harris

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 00:08:11