BBEdit操作求助:仅保留指定<tr>标签片段并删除其余内容
提取HTML中特定格式的
<tr>标签行 方法一:使用Grep命令(Linux/macOS/Windows WSL)
处理大文件时,Grep是高效的选择,直接运行以下命令:
grep -o '<tr id="post-[0-9]*"' your_file.html > output.txt
-o参数仅输出匹配到的内容片段,而非整行<tr id="post-[0-9]*"是匹配目标格式的正则表达式,[0-9]*匹配任意长度的数字后缀- 替换
your_file.html为你的源HTML文件路径,output.txt是保存结果的目标文件
如果使用Windows原生PowerShell,可执行:
Select-String -Path your_file.html -Pattern '<tr id="post-[0-9]*"' | ForEach-Object { $_.Matches.Value } | Out-File output.txt
方法二:使用Python脚本(跨平台)
适合需要自定义逻辑或无命令行工具的场景,脚本如下:
import re # 替换为实际文件路径 input_file = "your_file.html" output_file = "output.txt" # 定义匹配规则 match_pattern = re.compile(r'<tr id="post-\d+"') # 逐行读取大文件,避免内存溢出 with open(input_file, 'r', encoding='utf-8') as in_f, open(output_file, 'w', encoding='utf-8') as out_f: for line in in_f: matches = match_pattern.findall(line) for item in matches: out_f.write(f"{item}\n")
- 正则表达式
r'<tr id="post-\d+"'精准匹配目标格式,\d+匹配一个及以上数字 - 逐行处理文件,即使是超大体积的HTML也不会占用过多内存
- 每个匹配结果单独写入一行到输出文件
内容的提问来源于stack exchange,提问作者Manos Krokos
相关产品推荐
相关产品推荐

