You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Bash按序号对比文件夹,复制缺失文件至新文件夹

Bash Solution to Copy Missing Suffix-Matched Files

Alright, let's work through this problem efficiently—since you're dealing with 10 folders each holding 15k images, we need a solution that's fast and avoids unnecessary overhead. Here's how to do it with Bash, focusing on matching files by their numeric suffixes and copying only the missing ones to a new folder.

Assumptions First

I’m assuming your filenames follow a pattern where the numeric identifier is at the end (e.g., clip_0042.png, video_frame_12345.png). If your naming scheme is different, we can tweak the regex later—just adjust the pattern to fit your files!

Efficient Script for Batch Processing

This script handles all 10 folder pairs at once, uses fast in-memory lookups for large file sets, and safely handles filenames with spaces or special characters.

Version 1: Using Associative Arrays (Fastest for Large Datasets)

Requires Bash 4.0+ (most modern systems have this by default):

#!/bin/bash

# Enable nullglob to avoid issues when no .png files exist
shopt -s nullglob

# Loop through each of your 10 folder sets (adjust the path pattern to match your structure)
for set_dir in ./path/to/your/sets/set*; do
    # Define paths for each folder in the set
    originals="$set_dir/Originals"
    modified="$set_dir/Modified"
    missing="$set_dir/Missing_Frames"

    # Create the Missing_Frames folder if it doesn't exist
    mkdir -p "$missing"

    # Initialize an associative array to track existing suffixes in Modified
    declare -A existing_suffixes

    # Populate the array with all numeric suffixes from Modified files
    for file in "$modified"/*.png; do
        filename=$(basename "$file")
        # Extract the numeric suffix (adjust this regex to match your filename pattern!)
        # Current regex grabs digits after the last underscore before .png
        suffix=$(echo "$filename" | sed -E 's/.*_([0-9]+)\.png/\1/')
        existing_suffixes["$suffix"]=1
    done

    # Iterate through Originals and copy files with missing suffixes
    for file in "$originals"/*.png; do
        filename=$(basename "$file")
        suffix=$(echo "$filename" | sed -E 's/.*_([0-9]+)\.png/\1/')
        # Check if the suffix is NOT in our existing list
        if [[ -z ${existing_suffixes["$suffix"]} ]]; then
            cp "$file" "$missing/"
            echo "Copied missing file: $filename"
        fi
    done

    # Clean up the array for the next folder set
    unset existing_suffixes
done

Version 2: Using Temporary Files (Compatible with Older Bash Versions)

If you're stuck with Bash <4.0, this version uses a temporary file to store suffixes instead of an associative array:

#!/bin/bash

# Loop through each folder set
for set_dir in ./path/to/your/sets/set*; do
    originals="$set_dir/Originals"
    modified="$set_dir/Modified"
    missing="$set_dir/Missing_Frames"
    temp_file="$set_dir/modified_suffixes.txt"

    mkdir -p "$missing"

    # Extract all numeric suffixes from Modified and store in a sorted temp file
    find "$modified" -type f -name "*.png" -print0 | \
        xargs -0 basename | \
        sed -E 's/.*_([0-9]+)\.png/\1/' | \
        sort -n > "$temp_file"

    # Copy missing files from Originals
    find "$originals" -type f -name "*.png" -print0 | while IFS= read -r -d '' file; do
        filename=$(basename "$file")
        suffix=$(echo "$filename" | sed -E 's/.*_([0-9]+)\.png/\1/')
        # Check if suffix is not present in the temp file
        if ! grep -qxF "$suffix" "$temp_file"; then
            cp "$file" "$missing/"
            echo "Copied missing file: $filename"
        fi
    done

    # Delete the temporary file
    rm "$temp_file"
done

Key Tweaks for Your Filename Pattern

The critical part is the sed regex—adjust it to match how your numeric suffix is formatted:

  • If your files are named like frame1234.png (no underscore before digits): use sed -E 's/.*([0-9]+)\.png/\1/'
  • If your suffix is fixed-length (e.g., 4 digits): use sed -E 's/.*([0-9]{4})\.png/\1/' to avoid accidentally grabbing digits from other parts of the filename
  • If your files are just numbered like 1234.png: use sed -E 's/([0-9]+)\.png/\1/'

Performance Notes

  • The associative array version is significantly faster for large datasets (15k files) because lookups are O(1) instead of O(n) with grep.
  • Using find -print0 and read -d '' ensures we handle filenames with spaces, hyphens, or special characters without errors.

内容的提问来源于stack exchange,提问作者mfg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:27:24