You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

能否实现带增量校验和计算的文件复制?rsync -c效率问题咨询

Can we do file copy with incremental checksum calculation?

Absolutely, this is totally feasible—and it’s exactly the right approach when dealing with millions of files where rsync’s upfront checksum precomputation becomes a massive bottleneck. The key is to compute checksums on-the-fly as you read and copy the file data, instead of precomputing everything first. Let’s walk through practical solutions:

Command-Line Workarounds (Quick & Simple)

For single files or small batches, you can use standard Unix tools to pipe the file stream to both the destination and a checksum calculator simultaneously, using tee to split the stream:

# Copy a single file while computing its SHA256 checksum
cat /path/to/source/file | tee /path/to/target/file | sha256sum > /path/to/checksum.sha256

This way, you only read the source file once—no separate pass to compute the checksum first. For batch processing millions of files, wrap this in a loop with parallelization to avoid bottlenecks:

# Parallel batch processing (adjust -P to match your CPU cores)
find /source/dir -type f -print0 | xargs -0 -P 8 -I {} sh -c '
  target="/target/dir/${#source/dir#}{}"
  mkdir -p "$(dirname "$target")"
  cat "{}" | tee "$target" | sha256sum > "$target.sha256"
'

The -P 8 flag runs 8 concurrent copy/checksum processes—tweak this based on your system’s IO capacity to maximize speed without overwhelming disks.

Custom Scripts (Full Control)

If you need cross-platform support (e.g., Windows) or more logic (like verifying the checksum post-copy automatically), a simple Python script works wonders. Here’s a minimal example that reads files in chunks, writes to the target, and updates the checksum incrementally:

import hashlib
import os

def copy_with_checksum(source_path, target_path, chunk_size=1024*1024):
    # Create target directory if it doesn't exist
    os.makedirs(os.path.dirname(target_path), exist_ok=True)
    
    # Initialize checksum calculator
    sha256 = hashlib.sha256()
    
    with open(source_path, 'rb') as src, open(target_path, 'wb') as dst:
        while chunk := src.read(chunk_size):
            # Write chunk to target
            dst.write(chunk)
            # Update checksum with the chunk
            sha256.update(chunk)
    
    # Save the checksum
    checksum = sha256.hexdigest()
    with open(f"{target_path}.sha256", 'w') as f:
        f.write(f"{checksum}  {os.path.basename(target_path)}")
    
    # Optional: Verify target matches source checksum
    with open(target_path, 'rb') as dst:
        dst_sha256 = hashlib.sha256(dst.read()).hexdigest()
        if dst_sha256 != checksum:
            print(f"Warning: Checksum mismatch for {target_path}!")

# Example usage for a single file
copy_with_checksum("/source/file.txt", "/target/file.txt")

# Example batch processing
source_dir = "/source/dir"
target_dir = "/target/dir"
for root, _, files in os.walk(source_dir):
    for file in files:
        src_path = os.path.join(root, file)
        rel_path = os.path.relpath(src_path, source_dir)
        tgt_path = os.path.join(target_dir, rel_path)
        copy_with_checksum(src_path, tgt_path)

This script reads the source file in 1MB chunks (adjustable), writes each chunk to the target immediately, and updates the checksum incrementally—no upfront scanning or checksum precomputation. You can also add parallel processing here using concurrent.futures for faster batch operations.

A Note on rsync

If you still want to use rsync but avoid the upfront checksum hit, you can tweak its behavior:

  • Use --checksum-choice=auto to let rsync choose the most efficient checksum method.
  • Increase the block size with --block-size=1M to reduce the number of checksums computed (though this trades off precision for speed).
  • Avoid --archive if you don’t need metadata sync, as it adds overhead to the initial file list build.

That said, the on-the-fly copy+checksum methods above are still faster for large-scale datasets where precomputing all checksums is the main bottleneck.

内容的提问来源于stack exchange,提问作者Farshid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.08 18:12:27