You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用PowerShell从CSV文件列中提取<strong>标签内的值

Extract Values from <strong> Tags in a CSV Column with PowerShell

Got it, let's walk through how to pull out all the text wrapped in <strong> tags from your CSV column using PowerShell. Your sample data has HTML-encoded content (like &lt; instead of <), so we'll need to handle that first too.

Step 1: Break Down the Problem

Your CSV has a column with escaped HTML, and you need to grab every value inside <strong> tags. For your sample, that means extracting 487410-131, 502478-131, 555996-131, AEQT6T11110, and AEQT6T00210.

Step 2: Complete PowerShell Solution

Here's a script that reads your CSV, decodes the HTML, extracts the target values, and even lets you save the results back to a modified CSV:

# Replace with your actual CSV file path
$csvPath = "your-file.csv"
# Replace with the name of the column containing the HTML content
$targetColumn = "RelatedItems"

# Read the CSV into a PowerShell object collection
$csvData = Import-Csv -Path $csvPath -Delimiter ","

# Process each row to extract strong-tagged values
foreach ($row in $csvData) {
    # First, decode HTML entities (turn &lt; back into <, etc.)
    $decodedContent = [System.Web.HttpUtility]::HtmlDecode($row.$targetColumn)
    
    # Use regex to capture all text inside <strong> tags
    # The (.*?) is a non-greedy match to avoid grabbing content across multiple tags
    $extractedValues = [regex]::Matches($decodedContent, '<strong>(.*?)</strong>') | 
        ForEach-Object { $_.Groups[1].Value }
    
    # Optional: Print results for quick verification
    Write-Host "Row $($row[0]): Extracted values - $($extractedValues -join ', ')"
    
    # Optional: Add extracted values as a new column in the CSV
    $row | Add-Member -MemberType NoteProperty -Name "ExtractedStrongValues" -Value ($extractedValues -join '; ')
}

# Optional: Export modified CSV to a new file
$csvData | Export-Csv -Path "modified-file.csv" -NoTypeInformation -Encoding UTF8

Key Notes:

  • HTML Decoding: The [System.Web.HttpUtility]::HtmlDecode() method fixes escaped characters so our regex can properly match <strong> instead of &lt;strong&gt;.
  • Regex Logic: The pattern <strong>(.*?)</strong> uses non-greedy matching (.*?) to ensure we capture content from each individual <strong> tag, not from the first opening tag to the last closing tag in the entire string.
  • Flexibility: Adjust the delimiter in Import-Csv if your CSV uses tabs or another separator, and tweak the output logic to fit your needs (like saving values to a separate text file instead of a new CSV).

Testing with Your Sample Data

If you run this script with your exact sample content, you'll get this output:

Row 1: Extracted values - 487410-131, 502478-131, 555996-131, AEQT6T11110, AEQT6T00210

内容的提问来源于stack exchange,提问作者Jorge

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 08:13:44