寻求HTML转Contentful RichText的Python可行实现方案
实现HTML到Contentful RichText的Python方案
现有工具的局限性
- 官方仅提供反向转换(RichText转HTML)的
rich-text-renderer工具,无Python版的HTML转RichText方案 - 第三方npm包
contentful-html-rich-text-converter支持度有限,像<div>这类常见标签无法处理
可行的Python实现思路
1. 基于解析器自定义转换逻辑
用BeautifulSoup解析HTML,遍历DOM节点,手动映射到Contentful RichText的节点类型:
from bs4 import BeautifulSoup import json def html_to_rich_text(html): soup = BeautifulSoup(html, "html.parser") rich_text = {"data": {}, "content": [], "nodeType": "document"} def traverse(node): if node.name == "p": paragraph = {"data": {}, "content": [], "nodeType": "paragraph"} for child in node.children: if isinstance(child, str): paragraph["content"].append({ "data": {}, "content": child.strip(), "nodeType": "text" }) else: paragraph["content"].extend(traverse(child)) return [paragraph] # 可扩展支持更多标签,比如div转paragraph或嵌入节点 elif node.name == "div": div_content = {"data": {}, "content": [], "nodeType": "paragraph"} for child in node.children: if isinstance(child, str): div_content["content"].append({ "data": {}, "content": child.strip(), "nodeType": "text" }) else: div_content["content"].extend(traverse(child)) return [div_content] elif node.name == "strong": return [{ "data": {}, "content": traverse(node.contents)[0]["content"], "nodeType": "bold" }] # 继续添加h1-h6、list、blockquote等标签的映射逻辑 elif isinstance(node, str): if node.strip(): return [{ "data": {}, "content": node.strip(), "nodeType": "text" }] return [] return [] rich_text["content"] = traverse(soup) return json.dumps(rich_text, indent=2)
2. 扩展现有JS工具的Python调用
如果不想完全从零开发,可以通过subprocess调用contentful-html-rich-text-converter,同时自定义扩展它的标签支持:
- 先修改npm包的转换规则,添加
<div>等标签的处理逻辑 - 再用Python调用Node.js脚本完成转换
注意事项
- Contentful RichText有固定的节点类型规范,转换时需严格对应,避免无效节点
- 复杂HTML(如嵌套标签、自定义属性)需要针对性处理节点的
data字段
内容的提问来源于stack exchange,提问作者Eagleye
相关产品推荐
相关产品推荐

