使用PowerShell处理无表头CSV:保留指定列并去重
PowerShell Script to Filter Headerless CSV & Remove Duplicates
Got it, let's put together a PowerShell script that handles your headerless CSV exactly as you need—extracting columns 30, 31, and 39, then removing rows where the combination of these three columns is duplicated. Here's a robust, easy-to-adjust solution:
Basic Script (With Optional Headers in Output)
# Define your input and output file paths $inputCsv = "C:\path\to\your\source.csv" $outputCsv = "C:\path\to\your\filtered_output.csv" # Import the headerless CSV, assigning temporary column names (Col1, Col2, ..., Col100) # Adjust the 100 to match or exceed the total number of columns in your CSV $csvContent = Import-Csv -Path $inputCsv -Header (1..100 | ForEach-Object { "Col$_" }) # Extract the target columns and remove duplicates based on their combined values $filteredContent = $csvContent | Select-Object Col30, Col31, Col39 | Select-Object -Unique # Export the result to a new CSV $filteredContent | Export-Csv -Path $outputCsv -NoTypeInformation -UseQuotes AsNeeded
Key Details & Customizations
- Handling Non-Comma Delimiters: If your CSV uses a different delimiter (like
;or|), add the-Delimiterparameter toImport-Csv, e.g.,-Delimiter ';'. - Headerless Output: If you don't want the temporary
Col30/31/39headers in your output CSV, replace the finalExport-Csvline with this to skip the header row:$filteredContent | ConvertTo-Csv -NoTypeInformation -UseQuotes AsNeeded | Select-Object -Skip 1 | Set-Content -Path $outputCsv -Encoding UTF8 - Case-Insensitive Duplicate Removal: By default,
Select-Object -Uniqueis case-sensitive. To treat "ADValue" and "advalue" as duplicates while preserving the original case, useGroup-Objectinstead:$filteredContent = $csvContent | Group-Object -Property Col30, Col31, Col39 -CaseSensitive:$false | ForEach-Object { $_.Group[0] } | # Keep the first occurrence of each unique group Select-Object Col30, Col31, Col39
Why This Works
Import-Csvwith-Headersafely parses even complex CSV structures (like fields with quotes or line breaks) that manual string splitting would mess up.Select-Object -Uniqueautomatically checks for full matches across all three selected columns—only rows where all three values are identical get removed.
内容的提问来源于stack exchange,提问作者cfoster5
相关产品推荐
相关产品推荐

