You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何检测两文件全量差异并遇特殊字符重复对比输出至第三文件

问题核心

你需要的是把file2按--拆分成多个子组,每个子组分别和file1对比,输出file1中不存在于该子组的内容,但你原脚本直接用diff对比两个完整文件,只能得到全局差异,完全没处理分组逻辑,这就是问题根源。

改进实现脚本

以下是符合需求的Bash脚本,核心逻辑是逐个处理file2的分组,每个分组和file1求差集:

#!/bin/bash

file1="file1"
file2="file2"
output_file="Filee"

# 清空输出文件
> "$output_file"

current_group=""
while IFS= read -r line; do
    if [[ "$line" == "--" ]]; then
        # 处理当前分组:计算file1与当前组的差集
        if [[ -n "$current_group" ]]; then
            # comm -23 取file1有、当前组没有的行(需先排序)
            comm -23 <(sort "$file1") <(echo "$current_group" | grep -v "^[0-9]:" | sort) >> "$output_file"
            echo "--" >> "$output_file"
            current_group=""
        fi
    else
        # 记录分组内容,分组编号行直接写入结果
        if [[ "$line" =~ ^[0-9]+:$ ]]; then
            echo "$line" >> "$output_file"
        fi
        current_group+="$line"$'\n'
    fi
done < "$file2"

# 处理最后一个无--结尾的分组
if [[ -n "$current_group" ]]; then
    comm -23 <(sort "$file1") <(echo "$current_group" | grep -v "^[0-9]:" | sort) >> "$output_file"
fi

echo "对比完成,结果已保存到 $output_file"
关键逻辑说明
  • 分组拆分:遍历file2时,遇到--就触发当前分组的对比操作
  • 差集计算:用comm -23命令高效求差集,需要先对两个内容排序(comm要求输入有序),同时过滤掉分组的编号行(如1:)
  • 边界处理:单独处理file2末尾没有--的情况,避免最后一个分组被遗漏
测试结果

用你提供的示例文件运行脚本,输出文件内容与预期完全一致:

1:
a
c
--
2:
b
c
d
--
3:
b
d

内容的提问来源于stack exchange,提问作者Nadeen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.18 16:13:00