如何批量从HTML文件提取<title>并在<body>开头添加对应<h1>
批量处理遗留HTML文件:提取标题并添加
标签
我来帮你搞定这个批量处理的需求!针对遗留HTML文件里仅存于<title>元数据的标题,要在<body>开头插入对应的<h1>标签,这里有几个实用的方案,你可以根据自己的技术熟悉程度选择:
方案一:Python + BeautifulSoup(最可靠,推荐)
如果你的HTML文件结构可能不太规范(比如标签带空格、属性),用专业的HTML解析库比正则更稳妥,能避免匹配出错。
步骤:
- 先安装BeautifulSoup:
pip install beautifulsoup4 - 运行下面的脚本:
from bs4 import BeautifulSoup import os def add_h1_to_htmls(directory): # 遍历目标目录下的所有HTML文件 for filename in os.listdir(directory): if not filename.endswith(".html"): continue file_path = os.path.join(directory, filename) # 读取并解析HTML with open(file_path, 'r', encoding='utf-8') as f: soup = BeautifulSoup(f, 'html.parser') # 提取<title>内容 title_tag = soup.title if not title_tag or not title_tag.string: print(f"⚠️ 文件 {filename} 找不到有效<title>标签,跳过处理") continue title_text = title_tag.string.strip() # 找到<body>标签并插入<h1> body_tag = soup.body if not body_tag: print(f"⚠️ 文件 {filename} 找不到<body>标签,跳过处理") continue # 创建<h1>标签并插入到<body>最开头 h1_tag = soup.new_tag('h1') h1_tag.string = title_text body_tag.insert(0, h1_tag) # 加个换行让代码格式更整洁 body_tag.insert(1, '\n ') # 写回修改后的内容 with open(file_path, 'w', encoding='utf-8') as f: f.write(str(soup)) print(f"✅ 已处理完成:{filename}") # 替换成你的HTML文件所在目录路径 target_dir = "./your-html-folder" add_h1_to_htmls(target_dir)
注意事项:
- 批量处理前一定要先备份原文件,或者先拿1-2个测试文件验证效果
- 如果你的文件编码不是UTF-8,调整
encoding参数(比如gbk)
方案二:Python + 正则表达式(轻量,适合规范的HTML)
如果你的HTML文件结构很规范(<title>和<body>标签格式统一),用正则就能快速搞定:
import os import re def batch_add_h1(directory): for filename in os.listdir(directory): if not filename.endswith(".html"): continue file_path = os.path.join(directory, filename) with open(file_path, 'r', encoding='utf-8') as f: content = f.read() # 提取<title>内容(忽略大小写,兼容标签内的空格) title_match = re.search(r'<title\s*>(.*?)</title\s*>', content, re.IGNORECASE) if not title_match: print(f"⚠️ 文件 {filename} 无<title>标签,跳过") continue title_text = title_match.group(1).strip() # 在<body>标签后插入<h1>(兼容<body>带属性的情况,比如<body class="page">) updated_content = re.sub( r'(<body\b.*?>)', r'\1\n <h1>{}</h1>'.format(title_text), content, flags=re.IGNORECASE ) with open(file_path, 'w', encoding='utf-8') as f: f.write(updated_content) print(f"✅ 已处理完成:{filename}") # 替换成你的目标目录 batch_add_h1("./your-html-folder")
方案三:命令行工具(Linux/macOS:sed)
如果你习惯用命令行,sed可以快速批量处理,无需写脚本:
# 批量处理当前目录下的所有.html文件 # macOS用BSD sed,需要加-i '';Linux用GNU sed,直接写-i即可 for file in *.html; do sed -i '' -E ' # 捕获<title>内容到临时缓冲区 /<title>/{ h s/<title>(.*)<\/title>/\1/ x } # 找到<body>标签,插入<h1> /<body>/{ G s/<body>/<body>\n<h1>\1<\/h1>/ } ' "$file" done
方案四:Windows PowerShell
Windows用户可以用PowerShell完成批量处理:
# 替换成你的HTML文件目录 $targetDir = "./your-html-folder" Get-ChildItem -Path $targetDir -Filter "*.html" | ForEach-Object { $content = Get-Content $_.FullName -Raw -Encoding UTF8 # 提取<title>内容 $titleMatch = [regex]::Match($content, '<title\s*>(.*?)</title\s*>', [System.Text.RegularExpressions.RegexOptions]::IgnoreCase) if (-not $titleMatch.Success) { Write-Warning "⚠️ 文件 $($_.Name) 无有效<title>,跳过" return } $titleText = $titleMatch.Groups[1].Value.Trim() # 在<body>开头插入<h1> $updatedContent = [regex]::Replace($content, '(<body\b.*?>)', "`$1`n <h1>$titleText</h1>", [System.Text.RegularExpressions.RegexOptions]::IgnoreCase) # 写回文件 Set-Content $_.FullName -Value $updatedContent -Encoding UTF8 Write-Host "✅ 已处理完成:$($_.Name)" }
内容的提问来源于stack exchange,提问作者markwk
相关产品推荐
相关产品推荐

