You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

PowerShell提取HTML内容报错及批量处理需求

HTML文件指定标签内容批量提取脚本优化

初始问题

用户原本用PowerShell脚本提取单个HTML文件中指定标签范围内的内容,初始代码如下:

$search="<table id="
$linenumber= Get-Content ".\145039.html" | select-string $search | Select-Object LineNumber

$search="</table>"
$linenumber2= Get-Content ".\145039.html" | select-string $search | Select-Object LineNumber
#$linenumber2

# 需要提取的行号范围
$linesToFetch = $linenumber[2]..$linenumber2[2]
$currentLine  = 1
$result  = switch -File ".\145039.html" {
    default { if ($linesToFetch -contains $currentLine++) { $_ }}
}

# 写入文件并在控制台输出
$result | Set-Content -Path ".\excerpt.html" -PassThru

执行时触发类型转换错误:

Cannot convert the "@{LineNumber=6189}" value of type              "Selected.Microsoft.PowerShell.Commands.MatchInfo" to type "System.Int32".
At line:10 char:1
+ $linesToFetch = $linenumber[2]..$linenumber2[2]
+ ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
    + CategoryInfo          : InvalidArgument: (:) [], RuntimeException
    + FullyQualifiedErrorId : ConvertToFinalInvalidCastException

原因是$linenumber和$linenumber2返回的是包含LineNumber属性的对象,而非纯整数:

LineNumber
----------
      6015

单文件解决方案

通过添加-ExpandProperty LineNumber参数提取纯行号后,单文件提取功能恢复正常,更新后的代码如下:

$search1="disconnect-status"
$linenumber1= Get-Content ".\145039.html" | select-string $search1 | Select-Object -ExpandProperty LineNumber

$search2="</table>"
$linenumber2= Get-Content ".\145039.html" | select-string $search2 | Select-Object -ExpandProperty LineNumber

# 需要提取的行号范围
$linesToFetch = $linenumber1[3]..$linenumber2[1]
$currentLine  = 1
$result  = switch -File ".\145039.html" {
    default { if ($linesToFetch -contains $currentLine++) { $_ }}
}

# 写入文件并在控制台输出
$result | Set-Content -Path ".\excerpt.html" -PassThru

批量处理实现脚本

要遍历目录下所有HTML文件并批量提取指定范围内容,可使用以下脚本:

# 定义起始和结束搜索关键词
$startSearch = "disconnect-status"
$endSearch = "</table>"

# 获取当前目录下所有HTML文件
$htmlFiles = Get-ChildItem -Path .\ -Filter *.html

foreach ($file in $htmlFiles) {
    # 获取起始行号(对应原脚本的第4个匹配项,索引为3)
    $startLine = (Get-Content $file.FullName | Select-String $startSearch | Select-Object -ExpandProperty LineNumber)[3]
    # 获取结束行号(对应原脚本的第2个匹配项,索引为1)
    $endLine = (Get-Content $file.FullName | Select-String $endSearch | Select-Object -ExpandProperty LineNumber)[1]

    if ($startLine -and $endLine) {
        # 生成输出文件名,避免覆盖原文件
        $outputFile = Join-Path $file.DirectoryName ($file.BaseName + "_excerpt.html")
        
        $currentLine = 1
        $result = switch -File $file.FullName {
            default { 
                # 直接判断行号范围,大文件下性能更优
                if ($currentLine -ge $startLine -and $currentLine -le $endLine) {
                    $_
                }
                $currentLine++
            }
        }

        # 写入文件并在控制台输出结果
        $result | Set-Content -Path $outputFile -PassThru
        Write-Host "已处理文件:$($file.Name),输出路径:$outputFile"
    }
    else {
        Write-Warning "文件 $($file.Name) 未找到匹配的起始或结束标签,跳过处理"
    }
}

脚本说明

  • 自动遍历当前目录下所有.html文件,无需手动指定文件名
  • 每个文件生成独立的输出文件(格式为原文件名_excerpt.html),防止内容覆盖
  • 优化行范围判断逻辑,用-ge和-le替代-contains,处理大文件时效率更高
  • 增加异常判断,若文件未找到匹配标签则跳过并给出提示

内容的提问来源于stack exchange,提问作者SteveMize

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 10:21:04