Bash脚本优化:大文件多行复制效率提升需求
脚本优化方案
原脚本处理大文件慢的核心问题:循环中每次调用sed -i都会完整读取并重新写入目标文件,行数越多,重复IO的开销就越大,直接拖慢速度。下面是两种高效优化方案,均只对目标文件做一次修改操作:
方法一:临时文件中转(兼容性强)
这种方法适合所有支持sed的环境,步骤清晰:
#!/bin/bash source_file="..." dest_file="..." first_line_to_copy=... last_line_to_copy=... dest_line=... # 1. 提取源文件中需要复制的行到临时文件 sed -n "${first_line_to_copy},${last_line_to_copy}p" "$source_file" > /tmp/insert_content.tmp # 2. 一次性将临时文件内容插入到目标文件的指定位置(dest_line之前) # 逻辑:先输出目标文件前dest_line-1行,再输出插入内容,最后输出dest_line及以后的行 sed -e "1,${dest_line}-1p" -e "/^$/r /tmp/insert_content.tmp" -e "${dest_line},\$p" "$dest_file" > /tmp/new_dest.tmp # 3. 替换原目标文件并清理临时文件 mv /tmp/new_dest.tmp "$dest_file" rm /tmp/insert_content.tmp
方法二:用awk一次性完成(无临时文件)
借助awk可以直接在内存中处理插入逻辑,减少临时文件的IO操作:
#!/bin/bash source_file="..." dest_file="..." first_line_to_copy=... last_line_to_copy=... dest_line=... # 先把需要插入的内容读入变量 insert_content=$(sed -n "${first_line_to_copy},${last_line_to_copy}p" "$source_file") # awk处理:当行号等于dest_line时,先打印插入内容,再打印当前行 awk -v insert="$insert_content" -v target_line="$dest_line" ' NR == target_line { print insert } 1 ' "$dest_file" > /tmp/new_dest.tmp && mv /tmp/new_dest.tmp "$dest_file"
关键优化点说明
两种方案都只对目标文件执行一次读+一次写操作,彻底避免了原脚本中每行都修改文件的重复IO开销,处理大文件时速度会有数量级的提升。
内容的提问来源于stack exchange,提问作者Jean-Luc Delarbre
相关产品推荐
相关产品推荐

