使用PowerShell从CSV文件列中提取<strong>标签内的值
Got it, let's walk through how to pull out all the text wrapped in <strong> tags from your CSV column using PowerShell. Your sample data has HTML-encoded content (like < instead of <), so we'll need to handle that first too.
Step 1: Break Down the Problem
Your CSV has a column with escaped HTML, and you need to grab every value inside <strong> tags. For your sample, that means extracting 487410-131, 502478-131, 555996-131, AEQT6T11110, and AEQT6T00210.
Step 2: Complete PowerShell Solution
Here's a script that reads your CSV, decodes the HTML, extracts the target values, and even lets you save the results back to a modified CSV:
# Replace with your actual CSV file path $csvPath = "your-file.csv" # Replace with the name of the column containing the HTML content $targetColumn = "RelatedItems" # Read the CSV into a PowerShell object collection $csvData = Import-Csv -Path $csvPath -Delimiter "," # Process each row to extract strong-tagged values foreach ($row in $csvData) { # First, decode HTML entities (turn < back into <, etc.) $decodedContent = [System.Web.HttpUtility]::HtmlDecode($row.$targetColumn) # Use regex to capture all text inside <strong> tags # The (.*?) is a non-greedy match to avoid grabbing content across multiple tags $extractedValues = [regex]::Matches($decodedContent, '<strong>(.*?)</strong>') | ForEach-Object { $_.Groups[1].Value } # Optional: Print results for quick verification Write-Host "Row $($row[0]): Extracted values - $($extractedValues -join ', ')" # Optional: Add extracted values as a new column in the CSV $row | Add-Member -MemberType NoteProperty -Name "ExtractedStrongValues" -Value ($extractedValues -join '; ') } # Optional: Export modified CSV to a new file $csvData | Export-Csv -Path "modified-file.csv" -NoTypeInformation -Encoding UTF8
Key Notes:
- HTML Decoding: The
[System.Web.HttpUtility]::HtmlDecode()method fixes escaped characters so our regex can properly match<strong>instead of<strong>. - Regex Logic: The pattern
<strong>(.*?)</strong>uses non-greedy matching (.*?) to ensure we capture content from each individual<strong>tag, not from the first opening tag to the last closing tag in the entire string. - Flexibility: Adjust the delimiter in
Import-Csvif your CSV uses tabs or another separator, and tweak the output logic to fit your needs (like saving values to a separate text file instead of a new CSV).
Testing with Your Sample Data
If you run this script with your exact sample content, you'll get this output:
Row 1: Extracted values - 487410-131, 502478-131, 555996-131, AEQT6T11110, AEQT6T00210
内容的提问来源于stack exchange,提问作者Jorge
相关产品推荐
相关产品推荐

