Ubuntu 20下4GB大文件按行内容分类至指定文件的Bash问题
大文件按行分类:问题排查与高效解决方案
问题场景
在Ubuntu 20系统中处理4GB级大文件,需求是将文件中每行根据包含的特定字符串分类到对应文件,但编写的Bash循环脚本无法实现预期效果,耗时数小时仍未解决。
测试用例
测试文件 test.log
go to file 1 should go to file 2 should be seen in file 3 another for file 3 I belong in file 3 file 2 is my place
原错误脚本 processLog.sh
while read -r LINE do grep -h "file 1" > file1.txt grep -h "file 2" > file2.txt grep -h "file 3" > file3.txt done < test.log
预期输出结果
file1.txt go to file 1 file2.txt should go to file 2 file 2 is my place file3.txt should be seen in file 3 another for file 3 I belong in file 3
原脚本问题分析
- 逻辑错误:循环内的
grep命令未接收当前行$LINE的内容,默认等待标准输入,导致完全无法按行匹配分类; - 效率极低:即便修正逻辑,用
while read逐行处理搭配多次grep调用,对于4GB大文件会产生大量IO和进程创建开销,处理速度极慢。
高效解决方案
使用awk脚本仅遍历文件一次,即可完成分类,大幅提升处理速度,尤其适合大文件场景。最终采用的命令行调用脚本如下(用反斜杠实现换行):
awk '\ !/INSERT INTO / {print > "DATE_mysql_om.sql"; next } \ /INSERT INTO `objecttype`/ {print > "insert_om_objecttype.sql"; next } \ /INSERT INTO `objecttag`/ {print > "insert_om_objecttag.sql"; next } \ /INSERT INTO `object_objecttag`/ {print > "insert_om_object_objecttag.sql"; next } \ /INSERT INTO `indexlogger`/ {print > "insert_om_indexlogger.sql";}' \ *_mysqldump.sql
脚本逻辑说明
- 匹配不包含
INSERT INTO的行,写入DATE_mysql_om.sql; - 匹配含不同表名的
INSERT INTO语句,分别写入对应的SQL文件; - 全程仅遍历目标文件一次,处理4GB大文件效率极高。
内容的提问来源于stack exchange,提问作者user3008410
相关产品推荐
相关产品推荐

