使用PowerShell从文本文件提取跨换行URL并处理等号
提取并处理邮件中的浏览器查看URL
我有一个大型文本文件,里面包含“View this email in your browser”字样及后续的URL。URL格式不固定,有时会跨换行显示;当URL跨换行时,需要删除换行末尾的等号,但要保留URL中其他位置的等号。
示例文本如下:
View this email in your browser (https://us15.campaign-archive.com/?e=3D1460&u=3Df6e2bb1612577510b&id=3D2c8be)
View this email in your browser <https://mail.com/?e=3D14=
60&u=3Df612577510b&id=3D2c8be>View this email in your browser (https://eg.com/?e=3D1460&u=3Df6510b&id=3D2c8be)
我需要用PowerShell提取这些URL,去除可能存在的()或<>包裹符号,之后把URL对应的内容下载为HTML文件。目前已有部分代码片段:
if ($str -match '(?<=\()https?://[^)]+') { # # ... remove any line breaks from it, and output the result. $Matches.0 -replace '\r?\n' } if ($str -match '(?<=\<)https?://[^>]+') { # # ... remove any line breaks from it, and output the result. $Matches.0 -replace '\r?\n' }
完整实现方案
最终代码
# 替换为你的目标文本文件路径 $targetFile = "你的文件路径.txt" $fileContent = Get-Content -Path $targetFile -Raw # 匹配所有符合规则的URL(兼容跨换行场景) $urlPattern = '(?<=View this email in your browser\s*(?:\(|<))https?://[\s\S]+?(?=(?:\)|>))' $urlMatches = [regex]::Matches($fileContent, $urlPattern) foreach ($match in $urlMatches) { # 清理URL:删除换行前的等号 + 移除所有换行符 $cleanUrl = $match.Value -replace '=\r?\n', '' -replace '\r?\n', '' # 自定义输出文件名,这里用URL哈希值避免重复 $outputName = "$($cleanUrl.GetHashCode()).html" # 下载并保存HTML文件 Invoke-WebRequest -Uri $cleanUrl -OutFile $outputName Write-Host "已完成下载:$outputName" }
代码说明
- 读取文件:使用
Get-Content -Raw一次性读取整个文件内容,确保能捕捉到跨换行的URL。 - 正则匹配:
- 用
(?<=...)正向断言锁定目标URL的前置规则(必须跟在指定文本及(/<之后); - 用
[\s\S]+?匹配任意字符(包括换行),避免因URL跨换行导致匹配中断; - 用
(?=...)正向断言锁定URL的结束标记()/>)。
- 用
- URL清理:
- 先替换
=\r?\n为空,精准删除跨换行时末尾的等号; - 再移除剩余的换行符,确保URL是连续有效的字符串。
- 先替换
- 下载保存:调用
Invoke-WebRequest下载HTML内容,用URL的哈希值作为文件名,避免重复覆盖。
内容的提问来源于stack exchange,提问作者Coder
相关产品推荐
相关产品推荐

