Perl脚本问题:CSV文件按列去重并累加对应列数值
Hey there! Let's tackle this CSV task together—since you already know how to read and write files in Perl, we’ll focus on the core logic: tracking duplicates and accumulating their corresponding values.
The Core Idea: Use a Perl Hash
The perfect tool for this job is a Perl hash (associative array). We’ll use the value from your "duplicate check column" as the hash key (since hash keys are always unique), and the hash value will store the running total from the column you want to sum up.
Step-by-Step Example Code
Let’s assume your CSV has two columns: say, a Product column (where we check for duplicates) and a Quantity column (the values we want to sum). Here’s a complete script that does exactly what you need:
#!/usr/bin/perl use strict; use warnings; # Hash to store unique keys and their accumulated totals my %totals; # Open input CSV file open my $in_fh, '<', 'input.csv' or die "Can't open input.csv: $!"; # Skip the header row (remove this line if your CSV has no header) my $header = <$in_fh>; # Process each line in the input file while (my $line = <$in_fh>) { chomp $line; # Remove the trailing newline # Split the line into columns (adjust delimiter if needed, e.g., \t for tabs) my ($product, $quantity) = split /,/, $line; # Ensure the quantity is treated as a number (avoids string concatenation bugs) $quantity += 0; # Update the hash: if the product exists, add the quantity; else, initialize it if (exists $totals{$product}) { $totals{$product} += $quantity; } else { $totals{$product} = $quantity; } # Pro tip: You can shorten the above if/else to one line: # $totals{$product} += $quantity || 0; } close $in_fh; # Open output file to write results open my $out_fh, '>', 'output.csv' or die "Can't open output.csv: $!"; # Write header to output (customize this to match your columns) print $out_fh "Product,Total_Quantity\n"; # Loop through the hash and write each unique product with its total foreach my $product (sort keys %totals) { print $out_fh "$product,$totals{$product}\n"; } close $out_fh; print "Success! Results saved to output.csv\n";
Key Details to Note
use strict; use warnings;: Always include these—they’ll catch silly mistakes (like typos in variable names) that are easy for new Perl devs to miss.- Handling Complex CSVs: If your CSV has fields with commas (e.g.,
"Smith, Jane") or quoted values, the basicsplitmethod will break. For these cases, use theText::CSVmodule (install it viacpanm Text::CSV). Here’s a quick snippet using it:use Text::CSV; my $csv = Text::CSV->new({ binary => 1, auto_diag => 1 }); open my $in_fh, '<', 'input.csv' or die "Can't open input.csv: $!"; $csv->getline($in_fh); # Skip header while (my $row = $csv->getline($in_fh)) { my ($product, $quantity) = @$row; $totals{$product} += $quantity || 0; } - Sorting Results: The
sort keys %totalsline sorts your unique values alphabetically in the output—removesortif you don’t care about order.
That’s it! The hash takes care of automatically tracking unique values, and each time you encounter a duplicate key, you just add the new value to the existing total.
内容的提问来源于stack exchange,提问作者Chris Simmons

