Arch Linux下快速检测文件系统中是否存在内容相同但文件名不同的图片
Hey there! Dealing with old CD images and a huge existing library is definitely a hassle, but we can solve this efficiently with some command-line tools and a bit of scripting. Since you mentioned EXIF might differ but you care about the actual image binary data, we'll focus on hashing just the pixel content (ignoring metadata) to find duplicates.
First, let's get the necessary tools installed on your Arch system:
sudo pacman -S imagemagick perl-image-exiftool
Step 1: Precompute a Hash Database for Your Existing Images
Since you can narrow down to 20-30k images, we'll first generate a list of hashes based on their pixel data (stripping out all metadata). This is a one-time step, so take the time to do it right—parallel processing will speed this up.
Run this command, replacing /path/to/your/target/folders with the folders you want to check:
# Find all common image files and process them in parallel find /path/to/your/target/folders -type f \( -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.png" -o -iname "*.gif" -o -iname "*.bmp" \) | xargs -P 4 -I {} sh -c ' # Convert image to raw pixel data (stripping metadata) and compute MD5 hash hash=$(convert "{}" -strip -depth 8 rgb:- 2>/dev/null | md5sum | awk '\''{print $1}'\'') # Only add valid hashes to the database (skip corrupted files) if [ -n "$hash" ]; then echo "$hash {}" else echo "Error processing: {}" >> image_error_log.txt fi ' >> existing_image_hashes.txt
-P 4uses 4 CPU cores—adjust this number based on how many cores your system has (e.g.,-P 8for an 8-core CPU).convert -stripremoves all metadata (EXIF, etc.), so we're only hashing the actual pixel data.- We log any corrupted files to
image_error_log.txtso you can check them later.
Next, sort the hash file to make lookups much faster:
sort existing_image_hashes.txt > sorted_existing_hashes.txt
Step 2: Check CD Images Against the Hash Database
Now, for each CD, we'll mount it, compute hashes for its images, and compare them against our precomputed database.
- Mount the CD first:
sudo mount /dev/cdrom /mnt/cdrom
(If /mnt/cdrom doesn't exist, create it with sudo mkdir -p /mnt/cdrom.)
- Run this script to check for duplicates:
find /mnt/cdrom -type f \( -iname "*.jpg" -o -iname "*.jpeg" -o -iname "*.png" -o -iname "*.gif" -o -iname "*.bmp" \) | while read -r cd_img; do # Compute hash for the CD image hash=$(convert "$cd_img" -strip -depth 8 rgb:- 2>/dev/null | md5sum | awk '{print $1}') if [ -z "$hash" ]; then echo "Corrupted or unreadable image: $cd_img" >> cd_error_log.txt continue fi # Check if the hash exists in our sorted database if fgrep -q "$hash" sorted_existing_hashes.txt; then # Get the path to the existing duplicate existing_path=$(fgrep "$hash" sorted_existing_hashes.txt | awk '{print $2}') echo "✅ Duplicate found: $cd_img matches existing file at $existing_path" else echo "❌ No duplicate found for: $cd_img" # Optional: Uncomment below to copy new images to a folder # cp "$cd_img" /path/to/your/new/images/folder fi done
- When done, unmount the CD:
sudo umount /mnt/cdrom
Optimizations & Notes
- Hash Choice: We used MD5 because it's fast, but if you want a more collision-resistant hash, replace
md5sumwithsha1sumorsha256sum(they're slightly slower but still manageable). - Parallel Processing for CDs: If your CD drive is fast enough, you can use
xargs -Phere too, similar to Step 1, to speed up hash computation for CD images. - Perceptual Matching: If you want to find visually similar images (not exact pixel matches), you could use tools like
perceptualhash(from the AUR), but that's slower and not needed for your exact binary match requirement.
This approach should be fast enough for your needs—since we precompute the existing hashes once, each CD check will only process the images on the CD, which is way quicker than comparing every CD image against 20k files every time.
备注:内容来源于stack exchange,提问作者ductTapeIsMagic

