You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Perl脚本求助:按第二列唯一值筛选第三列数值不同的行

Perl Solution to Print Rows with Unique Third Column Values per Second Column

Let’s break down how to solve this problem. Based on your description, I assume you want to print one row for each unique combination of the second column value and third column value (i.e., for each group of rows sharing the same second column, you get one row per distinct third column entry). Here’s a straightforward Perl script that does exactly that:

#!/usr/bin/perl
use strict;
use warnings;

my %seen; # Tracks (col2, col3) pairs we've already printed

while (<>) {
    chomp;
    my @columns = split; # Split line into columns (whitespace-separated; adjust if needed)
    
    # Skip lines that don't have at least 3 columns
    next unless scalar @columns >= 3;
    
    my ($col2, $col3) = ($columns[1], $columns[2]);
    
    # If we haven't seen this (col2, col3) pair before, print the line and mark it as seen
    unless ($seen{$col2}{$col3}) {
        print "$_\n";
        $seen{$col2}{$col3} = 1;
    }
}

How This Works:

  • Reading Input: The <> operator reads from standard input or any files you pass as command-line arguments (e.g., ./script.pl input.txt).
  • Tracking Unique Pairs: The two-level hash %seen keeps track of which combinations of the second and third columns we’ve already printed. The first key is the second column value, and the inner key is the third column value.
  • Printing Unique Rows: For each line, we check if the (col2, col3) pair is new. If it is, we print the entire line and mark the pair as seen to avoid duplicates.

Adjustments for Different Delimiters:

If your input uses a delimiter other than whitespace (like commas for CSV), modify the split line. For example, for comma-separated values:

my @columns = split /,/;

Alternative: Print Rows Where the Second Column Group Has Multiple Distinct Third Columns

If you instead want to print all rows belonging to second column groups that have at least two different third column values (excluding groups where all third column values are the same), use this two-pass approach:

#!/usr/bin/perl
use strict;
use warnings;

my %col3_groups;

# First pass: Collect distinct third column values per second column
while (<>) {
    chomp;
    my @columns = split;
    next unless scalar @columns >=3;
    my ($col2, $col3) = ($columns[1], $columns[2]);
    $col3_groups{$col2}{$col3} = 1;
}

# Identify which second column groups have multiple distinct third columns
my %has_multiple_values;
foreach my $col2 (keys %col3_groups) {
    $has_multiple_values{$col2} = 1 if scalar keys %{$col3_groups{$col2}} > 1;
}

# Second pass: Re-read input and print rows from groups with multiple values
foreach my $file (@ARGV) {
    open my $fh, '<', $file or die "Couldn't open $file: $!";
    while (<$fh>) {
        chomp;
        my @columns = split;
        next unless scalar @columns >=3;
        my $col2 = $columns[1];
        print "$_\n" if $has_multiple_values{$col2};
    }
    close $fh;
}

This script first scans the input to find which second column groups have varying third column values, then re-reads the input to print all rows from those groups.

内容的提问来源于stack exchange,提问作者Sonal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 09:01:14