如何利用rsync高效追加大型文件?求参数解析与优化方案
大型.fastq.gz文件合并与高效传输问题
背景信息
我们公司经常需要合并5GB至200GB的大型.fastq.gz格式文件,这些文件存储在A~H号存储服务器上,通过sshfs挂载到本地/storage/archXX目录(archXX代表单台服务器的某块硬盘),合并后的最终文件需要存入生产服务器X。
原有追加方案及问题
此前我们使用以下命令实现文件追加:
pv /storage/archXX/customerID/fastq/data.fq.gz >> /work/customerID/fastq/data.fq.gz && pv /storage/archYY/customerID/fastq/data.fq.gz >> /work/customerID/fastq/data.fq.gz
该方法能保证.gz文件追加的数据完整性,但会给存储到生产服务器的传输带来极大系统负载,导致同事的数据访问速度变慢。
当前rsync中转方案及痛点
为降低负载,我们改用rsync直接从存储服务器复制文件至生产服务器,当前使用的脚本模板如下:
CustomerID=10 ARCH=arch49 #get the (ssh) path for the wanted archive ARCH_PATH=$(ssh production "df | grep -w ${ARCH} | sed 's/.*@//'" | awk '{print $1}') cd /work/${CustomerID}/fastq/ #copy directly to work/CustomerID/fastq/ rsync -r -v --progress user@${ARCH_PATH}/${CustomerID}/fastq/* /work/${CustomerID}/fastq/ ARCH=arch51 ARCH_PATH=$(ssh production "df | grep -w ${ARCH} | sed 's/.*@//'" | awk '{print $1}') cd /work/temp #copy to work/temp directory rsync -r -v --progress user@${ARCH_PATH}/${CustomerID}/fastq/* /work/temp/ #append files from /work/temp to /work/CustomerID/fastq/ pv "/work/temp/${CustomerID}_R1.fastq.gz" >> "/work/${CustomerID}/fastq/${CustomerID}_R1.fastq.gz" && pv "/work/temp/${CustomerID}_R2.fastq.gz" >> "/work/${CustomerID}/fastq/${CustomerID}_R2.fastq.gz"
此方案虽可行,但因需要在临时目录中转拷贝而效率低下,因此希望直接用rsync将文件追加至/work/CustomerID/fastq/data.gz。
关于rsync参数的疑问
查阅rsync手册时发现--append参数,手册说明该参数仅向较短的文件追加内容,但我无法确保服务器上的首个文件是最短的,且追加三个及以上文件时,先拷贝的文件可能比剩余文件更大。有以下疑问:
- 我对
--append参数的理解是否正确? - rsync是否有其他支持此类追加的参数?
- 该追加功能是否能如我预期工作?
其他高效传输方案询问
除了现有方案,还有哪些高效传输大型文件的网络方案?
内容的提问来源于stack exchange,提问作者Moritz
相关产品推荐
相关产品推荐

