能否实现带增量校验和计算的文件复制?rsync -c效率问题咨询
Absolutely, this is totally feasible—and it’s exactly the right approach when dealing with millions of files where rsync’s upfront checksum precomputation becomes a massive bottleneck. The key is to compute checksums on-the-fly as you read and copy the file data, instead of precomputing everything first. Let’s walk through practical solutions:
Command-Line Workarounds (Quick & Simple)
For single files or small batches, you can use standard Unix tools to pipe the file stream to both the destination and a checksum calculator simultaneously, using tee to split the stream:
# Copy a single file while computing its SHA256 checksum cat /path/to/source/file | tee /path/to/target/file | sha256sum > /path/to/checksum.sha256
This way, you only read the source file once—no separate pass to compute the checksum first. For batch processing millions of files, wrap this in a loop with parallelization to avoid bottlenecks:
# Parallel batch processing (adjust -P to match your CPU cores) find /source/dir -type f -print0 | xargs -0 -P 8 -I {} sh -c ' target="/target/dir/${#source/dir#}{}" mkdir -p "$(dirname "$target")" cat "{}" | tee "$target" | sha256sum > "$target.sha256" '
The -P 8 flag runs 8 concurrent copy/checksum processes—tweak this based on your system’s IO capacity to maximize speed without overwhelming disks.
Custom Scripts (Full Control)
If you need cross-platform support (e.g., Windows) or more logic (like verifying the checksum post-copy automatically), a simple Python script works wonders. Here’s a minimal example that reads files in chunks, writes to the target, and updates the checksum incrementally:
import hashlib import os def copy_with_checksum(source_path, target_path, chunk_size=1024*1024): # Create target directory if it doesn't exist os.makedirs(os.path.dirname(target_path), exist_ok=True) # Initialize checksum calculator sha256 = hashlib.sha256() with open(source_path, 'rb') as src, open(target_path, 'wb') as dst: while chunk := src.read(chunk_size): # Write chunk to target dst.write(chunk) # Update checksum with the chunk sha256.update(chunk) # Save the checksum checksum = sha256.hexdigest() with open(f"{target_path}.sha256", 'w') as f: f.write(f"{checksum} {os.path.basename(target_path)}") # Optional: Verify target matches source checksum with open(target_path, 'rb') as dst: dst_sha256 = hashlib.sha256(dst.read()).hexdigest() if dst_sha256 != checksum: print(f"Warning: Checksum mismatch for {target_path}!") # Example usage for a single file copy_with_checksum("/source/file.txt", "/target/file.txt") # Example batch processing source_dir = "/source/dir" target_dir = "/target/dir" for root, _, files in os.walk(source_dir): for file in files: src_path = os.path.join(root, file) rel_path = os.path.relpath(src_path, source_dir) tgt_path = os.path.join(target_dir, rel_path) copy_with_checksum(src_path, tgt_path)
This script reads the source file in 1MB chunks (adjustable), writes each chunk to the target immediately, and updates the checksum incrementally—no upfront scanning or checksum precomputation. You can also add parallel processing here using concurrent.futures for faster batch operations.
A Note on rsync
If you still want to use rsync but avoid the upfront checksum hit, you can tweak its behavior:
- Use
--checksum-choice=autoto let rsync choose the most efficient checksum method. - Increase the block size with
--block-size=1Mto reduce the number of checksums computed (though this trades off precision for speed). - Avoid
--archiveif you don’t need metadata sync, as it adds overhead to the initial file list build.
That said, the on-the-fly copy+checksum methods above are still faster for large-scale datasets where precomputing all checksums is the main bottleneck.
内容的提问来源于stack exchange,提问作者Farshid

