You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Grep无法匹配视觉相同字符串的问题排查与解决咨询

问题描述

我写了check_ff_prefs.sh脚本用来检查Firefox是否修改了我的偏好设置,结果遇到异常:视觉上完全相同的字符串,却无法被bash、grep、diff匹配。

脚本内容:

#!/bin/bash

check_match() {
    grep -qx -- "${1}" $2 
    exit_code=$?
    
    if [ "$exit_code" == "0" ]; then
        echo "does match $3"
    else
        echo "does not match $3"
    fi
}

FILE_1=/home/mcm/projects/dotfiles/firefox/prefs.js
FILE_2=/home/mcm/.mozilla/firefox/6zpaipqa.mcm/prefs.js

last_line_file_1=$(tail -n 1 $FILE_1)
last_line_file_2=$(tail -n 1 $FILE_2)

check_match "$last_line_file_1" "$FILE_1" "file1 -> file1"
check_match "$last_line_file_2" "$FILE_2" "file2 -> file2"

check_match "$last_line_file_1" "$FILE_2" "file1 -> file2"
check_match "$last_line_file_2" "$FILE_1" "file2 -> file1"
    
if [ "$last_line_file_1" == "$last_line_file_2" ]; then 
    echo "they match"
else
    echo "they don't match"
fi

echo $last_line_file_1 > last_line_file_1.txt
echo $last_line_file_2 > last_line_file_2.txt

diff last_line_file_1.txt last_line_file_2.txt

脚本输出:

does match file1 -> file1
does match file2 -> file2
does not match file1 -> file2
does not match file2 -> file1
they do not match
1c1
< user_pref("webgl.vendor-string-override", " ");
---> user_pref("webgl.vendor-string-override", " ");

我试过用line=${line//\"/\\\"}转义引号,但没用。现在想问两个问题:

  1. 这些行实际存在什么差异?
  2. 是否可以通过grep或其他工具忽略该差异?

解答

1. 实际差异分析

这两行的核心差异是空格的编码类型不同:其中一个是普通ASCII空格(U+0020),另一个是全角空格(U+3000)、不间断空格(U+00A0)这类非ASCII空白字符。它们视觉上完全一致,但底层编码不同,导致bash、grep、diff都判定为不同字符。

你可以用cat -A命令查看不可见字符来确认:

cat -A last_line_file_1.txt last_line_file_2.txt

普通ASCII空格会显示为空格,非ASCII空白会显示为M-bM-^@M-^这类转义序列,一眼就能区分差异。

2. 忽略差异的方法

方法一:统一转换空白字符后匹配

用tr命令把所有类型的空白字符转换成普通ASCII空格,再执行匹配逻辑:
修改check_match函数:

check_match() {
    # 将目标行和文件内容统一转换为普通空格
    target=$(echo "$1" | tr -s '[:space:]' ' ')
    grep -qx -- "$target" <(tr -s '[:space:]' ' ' "$2")
    exit_code=$?
    
    if [ "$exit_code" == "0" ]; then
        echo "does match $3"
    else
        echo "does not match $3"
    fi
}

tr -s '[:space:]' ' '会把所有空白字符(包括全角空格、不间断空格等)压缩并替换为普通空格。

方法二:grep正则匹配任意空白

把目标行中的空格替换为正则的空白字符类,让grep忽略空格类型差异:

check_match() {
    # 将目标行的空格替换为匹配任意空白的正则
    target=$(echo "$1" | sed 's/ /[[:space:]]/g')
    grep -qx -- "$target" "$2"
    exit_code=$?
    
    if [ "$exit_code" == "0" ]; then
        echo "does match $3"
    else
        echo "does not match $3"
    fi
}

方法三:diff直接忽略空白差异

如果只是比较文件内容,用diff的-w参数即可忽略所有空白字符的差异:

diff -w last_line_file_1.txt last_line_file_2.txt

内容的提问来源于stack exchange,提问作者robbmj

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.18 19:30:47