如何用Python将HTML转换为JSON?求高效简便解决方案
Python 实现 HTML 转 JSON 的简便方案
以下是两种实用的方案,无需自行从零构建转换逻辑,适合处理大型HTML文件:
方案一:BeautifulSoup + 自定义递归转换
利用BeautifulSoup的HTML解析能力,配合递归遍历生成JSON结构,灵活性高且适合处理大文件。
步骤:
- 安装依赖:
pip install beautifulsoup4 lxml
(lxml作为解析器,比默认的html.parser效率更高,适合大型文件)
- 转换代码:
from bs4 import BeautifulSoup import json def parse_element(element): node = {"tag": element.name} # 保留标签属性 if element.attrs: node["attributes"] = element.attrs # 提取非空白文本 text = element.get_text(strip=True) if text: node["text"] = text # 递归处理子标签 children = [] for child in element.children: if child.name: # 仅处理标签节点,跳过纯文本/注释等非标签内容 children.append(parse_element(child)) if children: node["children"] = children return node # 流式读取大型HTML文件,避免内存溢出 with open("large_html_file.html", "r", encoding="utf-8") as html_file: soup = BeautifulSoup(html_file, "lxml") # 可选择解析整个HTML根节点,或指定特定节点(如soup.body) json_result = parse_element(soup.html) # 写入JSON文件 with open("output.json", "w", encoding="utf-8") as json_file: json.dump(json_result, json_file, indent=2, ensure_ascii=False)
优势:
- 可根据需求自定义JSON结构(比如过滤特定标签、保留/忽略属性)
- 流式解析方式适合GB级别的大型HTML文件,内存占用可控
方案二:使用现成转换库html2json
如果不需要高度自定义,直接用封装好的库快速实现转换,代码更简洁。
步骤:
- 安装依赖:
pip install html2json
- 转换代码:
from html2json import convert import json # 读取HTML内容(超大型文件可分块读取后拼接,或结合生成器处理) with open("large_html_file.html", "r", encoding="utf-8") as html_file: html_content = html_file.read() # 一键转换为JSON结构 json_data = convert(html_content) # 保存结果 with open("output.json", "w", encoding="utf-8") as json_file: json.dump(json_data, json_file, indent=2, ensure_ascii=False)
注意:
- 该库默认会保留所有标签和属性,适合通用场景
- 若文件过大导致内存不足,建议采用方案一的流式解析方式
内容的提问来源于stack exchange,提问作者adamnitz
相关产品推荐
相关产品推荐

