You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何删除重复文本行并忽略行尾附带的日期部分?

文本列表按前缀去重实现方案

以下提供三种不同场景的实现方式,均以:::分隔符前的字符串作为去重判断依据,默认保留重复条目中首次出现的完整行。

方法1:Linux/macOS 命令行快速实现

一行awk命令即可完成操作,无需额外安装依赖:

awk -F ':::' '!seen[$1]++' your_input_file.txt > deduplicated_output.txt
  • -F ':::'指定行分割符为:::,$1自动对应:::前的路径部分
  • 用数组seen记录每个路径的出现次数,首次出现时判定为真,输出完整行,重复条目直接跳过

方法2:Python 跨平台实现

适用于需要兼容Windows/macOS/Linux多系统的场景,代码如下:

seen = set()
# 替换成你的输入、输出文件路径
with open("your_input_file.txt", "r", encoding="utf-8") as f_in, open("deduplicated_output.txt", "w", encoding="utf-8") as f_out:
    for line in f_in:
        raw_line = line.rstrip("\n")
        # 跳过空行
        if not raw_line:
            continue
        # 仅拆分第一个出现的:::,避免路径本身含特殊字符导致拆分错误
        prefix = raw_line.split(":::", 1)[0].strip()
        if prefix not in seen:
            seen.add(prefix)
            f_out.write(raw_line + "\n")

方法3:Windows PowerShell 实现

Windows系统无需额外装软件,直接运行PowerShell命令即可:

$seen = @{}
Get-Content .\your_input_file.txt | ForEach-Object {
    $prefix = $_.Split(':::', 2)[0].Trim()
    if (-not $seen.ContainsKey($prefix)) {
        $seen[$prefix] = $_
        $_
    }
} | Out-File .\deduplicated_output.txt -Encoding UTF8

若需要保留重复条目中最后一次出现的完整行,可调整对应逻辑:

  • awk 调整为awk -F ':::' '{line[$1]=$0} END{for(i in line) print line[i]}' your_input_file.txt > deduplicated_output.txt
  • Python 改为用字典存储,每次遇到相同前缀就覆盖值,最后遍历字典值写入文件即可

内容的提问来源于stack exchange,提问作者Michael Joseph Jackson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 10:39:00