在Zsh中移除尾随空白时触发致命错误:堆内存不足,求解决方案
Zsh脚本移除文件尾随空白时触发“堆内存不足”错误的原因与解决办法
问题描述
在使用Zsh脚本批量移除文件中的尾随空白时,出现了fatal error: out of heap memory错误,脚本代码如下:
#!/usr/bin/env zsh setopt nullglob setopt extendedglob for file in **/*; do if file -b $file | grep -q "text"; then raw=$(<$file) # Remove trailing whitespace raw=${raw//[[:blank:]]##$' '/$' '} # Other stuff like removing CR+LF works fine # raw=${raw//$' '} echo $file printf "%s " $raw > $file fi done
错误原因
- 一次性加载大文件到内存:脚本用
raw=$(<$file)把整个文件内容读入变量,遇到几十MB甚至更大的文本文件时,Zsh会将全部内容占用堆内存,直接触发内存不足。 - 字符串替换的高内存开销:Zsh的参数扩展
${raw//.../...}在处理超大字符串时,会生成原字符串的副本进行匹配替换,内存消耗翻倍,进一步加剧内存紧张。
解决办法
1. 使用sed流式处理(最推荐)
sed是专门的文本处理工具,采用逐行流式处理,不需要加载整个文件到内存,内存占用极低,修改后的脚本:
#!/usr/bin/env zsh setopt nullglob setopt extendedglob for file in **/*; do if file -b "$file" | grep -q "text"; then # 原地修改文件,移除每行末尾的空白字符 # macOS用sed -i'',Linux用sed -i sed -i'' -E 's/[[:blank:]]+$//' "$file" echo "$file" fi done
注:s/[[:blank:]]+$//表示匹配每行末尾的所有空白字符(空格、制表符)并删除。
2. 优化Zsh脚本为逐行处理
如果必须用Zsh处理,避免一次性读入大文件,改成逐行读取并处理:
#!/usr/bin/env zsh setopt nullglob setopt extendedglob for file in **/*; do if file -b "$file" | grep -q "text"; then # 创建临时文件存储处理结果 temp_file=$(mktemp) # 逐行读取文件,保留行首空白 while IFS= read -r line; do # 移除当前行末尾的所有空白 cleaned_line=${line%%[[:blank:]]##} echo "$cleaned_line" >> "$temp_file" done < "$file" # 用临时文件替换原文件 mv "$temp_file" "$file" echo "$file" fi done
3. 跳过超大文件
如果不需要处理大文本文件,可以在脚本中增加文件大小判断,跳过超过阈值的文件:
#!/usr/bin/env zsh setopt nullglob setopt extendedglob # 设置最大处理文件大小,这里设为100MB max_size=$((100 * 1024 * 1024)) for file in **/*; do if [[ -f "$file" ]]; then # macOS获取文件大小的方式,Linux替换为stat -c "%s" "$file" file_size=$(stat -f "%z" "$file") if (( file_size > max_size )); then echo "跳过超大文件: $file" continue fi if file -b "$file" | grep -q "text"; then sed -i'' -E 's/[[:blank:]]+$//' "$file" echo "$file" fi fi done
内容的提问来源于stack exchange,提问作者Mark Boulder
相关产品推荐
相关产品推荐

