You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PowerShell实现提取标签间值生成无空格分号分隔表格

问题:处理包含XML结构的文本文件,替换指定行内容

我有一个文本文件,前7行需保留原样,第8行包含<LineItemTable>结构,每个<LINEITEM>下有4个<LINEITEMFIELD>值(表格长度可变但每行固定4个值)。需提取这些值并去除空格,以分号分隔为每行一组,替换原第8行。我编写了PowerShell代码但无法得到期望输出,请求帮助。

输入文件内容

0001117945
14102022
0001056.98
GBP
0000000.00
0000000.00
\\GLORSAWA01\EHIShared\Remittance\UK01\UKI_REM_COL58652cbc13ca49aabf000.pdf
<LineItemTable><LINEITEM><LINEITEMFIELD>20220916</LINEITEMFIELD><LINEITEMFIELD>2525636         </LINEITEMFIELD><LINEITEMFIELD>0.00                            </LINEITEMFIELD><LINEITEMFIELD>246.05                          </LINEITEMFIELD></LINEITEM><LINEITEM><LINEITEMFIELD>20220920</LINEITEMFIELD><LINEITEMFIELD>2527541         </LINEITEMFIELD><LINEITEMFIELD>0.00                            </LINEITEMFIELD><LINEITEMFIELD>450.12                          </LINEITEMFIELD></LINEITEM><LINEITEM><LINEITEMFIELD>20220922</LINEITEMFIELD><LINEITEMFIELD>2531147         </LINEITEMFIELD><LINEITEMFIELD>0.00                            </LINEITEMFIELD><LINEITEMFIELD>360.81                          </LINEITEMFIELD></LINEITEM></LineItemTable>

期望输出文件

0001117945
14102022
0001056.98
GBP
0000000.00
0000000.00
\\GLORSAWA01\EHIShared\Remittance\UK01\UKI_REM_COL58652cbc13ca49aabf000.pdf
20220916;2525636;0.00;246.05
20220920;2527541;0.00;450.12
20220922;2531147;0.00;360.81

我编写的代码

$importfolder = "\.\PowerShell_script\"
$outputfolder = "\.\PowerShell_script\Output\"
$files = "\.\PowerShell script\*.txt"
$list = Get-ChildItem -Path $files | select Name


$find1 = "&lt;LineItemTable&gt;&lt;LINEITEM&gt;&lt;LINEITEMFIELD&gt;"
$find2 = "&lt;/LINEITEMFIELD&gt;&lt;LINEITEMFIELD&gt;"
$find3 = "&lt;/LINEITEMFIELD&gt;&lt;/LINEITEM&gt;&lt;LINEITEM&gt;&lt;LINEITEMFIELD&gt;"

$replace1 = ""
$replace2 = "`t"
$replace3 = "`n"

ForEach($file in $list){
    echo $file.Name
    $filename = $file.Name
    $file = $importfolder + $file.Name
    $outputfile = $outputfolder + $filename

    $filecontent = Get-Content $file | 
    ForEach-Object { 
        if($_ -Match $find1)
            {$_ -replace $replace1}         

        if($_ -Match $find2)
            {$_ -replace $replace2} 

        if($_ -Match $find3)
            {$_ -replace $replace3} 

        else {$_} # output the line as is
     } | Set-Content $outputfile
}

解决方案

原代码的核心问题在于:替换逻辑碎片化,无法完整提取并重组字段内容,也没有处理字段内的多余空格,同时路径拼接方式容易出错。以下是修正后的代码:

$importFolder = ".\PowerShell_script\"
$outputFolder = ".\PowerShell_script\Output\"

# 确保输出目录存在
if (-not (Test-Path $outputFolder)) {
    New-Item -ItemType Directory -Path $outputFolder | Out-Null
}

Get-ChildItem -Path (Join-Path $importFolder "*.txt") | ForEach-Object {
    Write-Host "处理文件: $($_.Name)"
    $outputFile = Join-Path $outputFolder $_.Name

    # 一次性读取整个文件内容,方便处理单行XML结构
    $allContent = Get-Content -Path $_.FullName -Raw -Encoding UTF8

    # 拆分出前7行表头和包含LineItem的行
    $contentParts = $allContent -split "`n", 8
    $headerLines = $contentParts[0..6]
    $lineItemContent = $contentParts[7]

    # 提取所有LINEITEMFIELD内的内容,同时去除前后空格
    $fieldMatches = [regex]::Matches($lineItemContent, '(?<=<LINEITEMFIELD>)\s*(.*?)\s*(?=</LINEITEMFIELD>)')
    $cleanedFields = $fieldMatches | ForEach-Object { $_.Groups[1].Value }

    # 每4个字段组成一行,用分号分隔
    $processedItems = for ($i = 0; $i -lt $cleanedFields.Count; $i += 4) {
        $cleanedFields[$i..($i + 3)] -join ';'
    }

    # 组合表头和处理后的行,写入输出文件
    ($headerLines + $processedItems) -join "`n" | Set-Content -Path $outputFile -Encoding UTF8
}

代码关键点说明

  1. 文件读取方式:使用Get-Content -Raw一次性读取整个文件,避免逐行处理时的XML结构拆分问题。
  2. 正则提取内容:用正向/反向断言精准提取<LINEITEMFIELD>标签内的文本,同时通过\s*匹配并去除字段前后的多余空格。
  3. 字段分组重组:通过循环每4个字段为一组,用分号连接成新行,完美匹配需求的格式。
  4. 路径处理:使用Join-Path拼接文件路径,避免手动拼接时的斜杠错误;同时提前创建输出目录,防止写入失败。
  5. 编码一致性:指定UTF8编码读写文件,避免乱码问题。

内容的提问来源于stack exchange,提问作者Jana

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 11:05:26