You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PowerShell技术问询:删除含重复单词的行及保留同特征组最后一行

PowerShell解决方案:处理重复单词行与分组保留最后一行

Let's break down your two requirements and fix the issues with your initial approach—sort | Get-Unique only works for exact consecutive duplicate lines, which doesn't cover either of your needs. Here's how to tackle each problem, and combine them into one pipeline:

需求1:删除包含重复单词的行

To filter out lines where any word appears more than once, we can split each line into individual words, then compare the count of original words vs. unique words. If they match, the line has no duplicates:

Get-Content 1.txt | Where-Object {
    # Split the line into words, ignoring empty spaces
    $words = $_.Split(' ', [System.StringSplitOptions]::RemoveEmptyEntries)
    # Only keep lines where all words are unique
    $words.Count -eq ($words | Select-Object -Unique).Count
}

需求2:按特征字符串分组,仅保留每组最后一行

For grouping by the feature string (the first element in each line, like a or b in your example) and keeping only the last entry per group, use Group-Object to cluster lines by their feature, then extract the last line from each group:

Get-Content 1.txt | Group-Object -Property { $_.Split(' ')[0] } | ForEach-Object {
    # Grab the last line in the grouped set
    $_.Group[-1]
}

合并两个需求(完整命令)

If you need to apply both filters together (first remove lines with duplicate words, then keep only the last line per feature group), chain the commands into a single pipeline:

Get-Content 1.txt | 
    # Step 1: Filter out lines with duplicate words
    Where-Object {
        $words = $_.Split(' ', [System.StringSplitOptions]::RemoveEmptyEntries)
        $words.Count -eq ($words | Select-Object -Unique).Count
    } |
    # Step 2: Group by feature string and keep last line per group
    Group-Object -Property { $_.Split(' ')[0] } |
    ForEach-Object { $_.Group[-1] } |
    # Optional: Save the cleaned output to a new file
    Out-File -FilePath cleaned_1.txt -Encoding UTF8

为什么你的初始命令失败了?

sort | Get-Unique works by sorting lines and removing exact consecutive duplicates. But your use case needs:

  • Detection of duplicate words within a single line (not duplicate lines)
  • Grouping by a shared feature string (not exact line matches)
  • Retaining the last entry per group (not just removing duplicates)

That's why this approach didn't meet your needs.

内容的提问来源于stack exchange,提问作者Bogdan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 04:29:20