PowerShell提取HTML内容报错及批量处理需求
HTML文件指定标签内容批量提取脚本优化
初始问题
用户原本用PowerShell脚本提取单个HTML文件中指定标签范围内的内容,初始代码如下:
$search="<table id=" $linenumber= Get-Content ".\145039.html" | select-string $search | Select-Object LineNumber $search="</table>" $linenumber2= Get-Content ".\145039.html" | select-string $search | Select-Object LineNumber #$linenumber2 # 需要提取的行号范围 $linesToFetch = $linenumber[2]..$linenumber2[2] $currentLine = 1 $result = switch -File ".\145039.html" { default { if ($linesToFetch -contains $currentLine++) { $_ }} } # 写入文件并在控制台输出 $result | Set-Content -Path ".\excerpt.html" -PassThru
执行时触发类型转换错误:
Cannot convert the "@{LineNumber=6189}" value of type "Selected.Microsoft.PowerShell.Commands.MatchInfo" to type "System.Int32". At line:10 char:1 + $linesToFetch = $linenumber[2]..$linenumber2[2] + ~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + CategoryInfo : InvalidArgument: (:) [], RuntimeException + FullyQualifiedErrorId : ConvertToFinalInvalidCastException
原因是$linenumber和$linenumber2返回的是包含LineNumber属性的对象,而非纯整数:
LineNumber ---------- 6015
单文件解决方案
通过添加-ExpandProperty LineNumber参数提取纯行号后,单文件提取功能恢复正常,更新后的代码如下:
$search1="disconnect-status" $linenumber1= Get-Content ".\145039.html" | select-string $search1 | Select-Object -ExpandProperty LineNumber $search2="</table>" $linenumber2= Get-Content ".\145039.html" | select-string $search2 | Select-Object -ExpandProperty LineNumber # 需要提取的行号范围 $linesToFetch = $linenumber1[3]..$linenumber2[1] $currentLine = 1 $result = switch -File ".\145039.html" { default { if ($linesToFetch -contains $currentLine++) { $_ }} } # 写入文件并在控制台输出 $result | Set-Content -Path ".\excerpt.html" -PassThru
批量处理实现脚本
要遍历目录下所有HTML文件并批量提取指定范围内容,可使用以下脚本:
# 定义起始和结束搜索关键词 $startSearch = "disconnect-status" $endSearch = "</table>" # 获取当前目录下所有HTML文件 $htmlFiles = Get-ChildItem -Path .\ -Filter *.html foreach ($file in $htmlFiles) { # 获取起始行号(对应原脚本的第4个匹配项,索引为3) $startLine = (Get-Content $file.FullName | Select-String $startSearch | Select-Object -ExpandProperty LineNumber)[3] # 获取结束行号(对应原脚本的第2个匹配项,索引为1) $endLine = (Get-Content $file.FullName | Select-String $endSearch | Select-Object -ExpandProperty LineNumber)[1] if ($startLine -and $endLine) { # 生成输出文件名,避免覆盖原文件 $outputFile = Join-Path $file.DirectoryName ($file.BaseName + "_excerpt.html") $currentLine = 1 $result = switch -File $file.FullName { default { # 直接判断行号范围,大文件下性能更优 if ($currentLine -ge $startLine -and $currentLine -le $endLine) { $_ } $currentLine++ } } # 写入文件并在控制台输出结果 $result | Set-Content -Path $outputFile -PassThru Write-Host "已处理文件:$($file.Name),输出路径:$outputFile" } else { Write-Warning "文件 $($file.Name) 未找到匹配的起始或结束标签,跳过处理" } }
脚本说明
- 自动遍历当前目录下所有
.html文件,无需手动指定文件名 - 每个文件生成独立的输出文件(格式为
原文件名_excerpt.html),防止内容覆盖 - 优化行范围判断逻辑,用
-ge和-le替代-contains,处理大文件时效率更高 - 增加异常判断,若文件未找到匹配标签则跳过并给出提示
内容的提问来源于stack exchange,提问作者SteveMize
相关产品推荐
相关产品推荐

