You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python如何提取HTML的title标签内容作为文件名批量保存文件

批量按标签内容命名保存HTML文件实现方案 <a class="header-anchor" href="#批量按标签内容命名保存html文件实现方案" aria-hidden="true">#</a></h1> <h2 id="前置依赖">前置依赖 <a class="header-anchor" href="#前置依赖" aria-hidden="true">#</a></h2> <p>仅需要用到Python标准库的<code>re</code>和<code>os</code>模块,无需安装额外第三方包。</p> <h2 id="完整可运行代码">完整可运行代码 <a class="header-anchor" href="#完整可运行代码" aria-hidden="true">#</a></h2> <pre class="hljs"><code class="language-python volc-pre-code">import re import os # 配置项:根据你自己的本地路径修改 ORIGIN_HTML_DIR = "./origin_html" # 存放500个原始HTML文件的目录 SAVE_NEW_HTML_DIR = "./new_html" # 处理后新HTML文件的保存目录 # 自动创建保存目录(如果不存在不会报错) os.makedirs(SAVE_NEW_HTML_DIR, exist_ok=True) # 遍历原始目录下所有文件 for filename in os.listdir(ORIGIN_HTML_DIR): # 只处理后缀为.html的文件,跳过其他类型文件 if not filename.lower().endswith(".html"): continue file_path = os.path.join(ORIGIN_HTML_DIR, filename) # 读取原始HTML文件完整内容 with open(file_path, "r", encoding="utf-8", errors="ignore") as f: html_content = f.read() # 提取title标签文本内容 title_content = "" try: title_match = re.search(r'<title[^>]*>(.*?)</title>', html_content, re.IGNORECASE) if title_match: title_content = title_match.group(1).strip() except Exception: pass # 按照规则处理生成合法文件名 if title_content: # 1. 空格替换为连字符 new_filename = title_content.replace(" ", "-") # 2. 统一转为全小写(和你给出的示例规则对齐,如需保留原大小写可注释这行) new_filename = new_filename.lower() # 3. 移除文件名中不允许的特殊字符,避免保存报错 new_filename = re.sub(r'[\\/*?:"<>|]', "", new_filename) # 4. 补全.html后缀 new_filename = f"{new_filename}.html" else: # 没有提取到title的文件做兜底处理,避免程序中断 new_filename = f"no_title_{filename}" # 保存新HTML文件到指定目录 new_file_path = os.path.join(SAVE_NEW_HTML_DIR, new_filename) with open(new_file_path, "w", encoding="utf-8") as f: f.write(html_content) </code></pre> <h2 id="代码说明">代码说明 <a class="header-anchor" href="#代码说明" aria-hidden="true">#</a></h2> <ul> <li>正则匹配做了兼容性优化,支持<code><TITLE></code>大写标签场景,非贪婪匹配避免标签嵌套导致提取错误</li> <li>自动过滤文件名非法字符,兼容Windows、Linux等不同操作系统的文件名规则</li> <li>无title标签的文件做了兜底处理,不会出现批量处理中途中断的问题</li> <li>保留原始HTML的全部内容,仅修改输出文件名</li> </ul> <p>内容的提问来源于stack exchange,提问作者user15950808</p>
相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 18:06:01