You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

寻求HTML转Contentful RichText的Python可行实现方案

实现HTML到Contentful RichText的Python方案

现有工具的局限性

  • 官方仅提供反向转换(RichText转HTML)的rich-text-renderer工具,无Python版的HTML转RichText方案
  • 第三方npm包contentful-html-rich-text-converter支持度有限,像<div>这类常见标签无法处理

可行的Python实现思路

1. 基于解析器自定义转换逻辑

用BeautifulSoup解析HTML,遍历DOM节点,手动映射到Contentful RichText的节点类型:

from bs4 import BeautifulSoup
import json

def html_to_rich_text(html):
    soup = BeautifulSoup(html, "html.parser")
    rich_text = {"data": {}, "content": [], "nodeType": "document"}
    
    def traverse(node):
        if node.name == "p":
            paragraph = {"data": {}, "content": [], "nodeType": "paragraph"}
            for child in node.children:
                if isinstance(child, str):
                    paragraph["content"].append({
                        "data": {},
                        "content": child.strip(),
                        "nodeType": "text"
                    })
                else:
                    paragraph["content"].extend(traverse(child))
            return [paragraph]
        # 可扩展支持更多标签,比如div转paragraph或嵌入节点
        elif node.name == "div":
            div_content = {"data": {}, "content": [], "nodeType": "paragraph"}
            for child in node.children:
                if isinstance(child, str):
                    div_content["content"].append({
                        "data": {},
                        "content": child.strip(),
                        "nodeType": "text"
                    })
                else:
                    div_content["content"].extend(traverse(child))
            return [div_content]
        elif node.name == "strong":
            return [{
                "data": {},
                "content": traverse(node.contents)[0]["content"],
                "nodeType": "bold"
            }]
        # 继续添加h1-h6、list、blockquote等标签的映射逻辑
        elif isinstance(node, str):
            if node.strip():
                return [{
                    "data": {},
                    "content": node.strip(),
                    "nodeType": "text"
                }]
            return []
        return []
    
    rich_text["content"] = traverse(soup)
    return json.dumps(rich_text, indent=2)

2. 扩展现有JS工具的Python调用

如果不想完全从零开发,可以通过subprocess调用contentful-html-rich-text-converter,同时自定义扩展它的标签支持:

  • 先修改npm包的转换规则,添加<div>等标签的处理逻辑
  • 再用Python调用Node.js脚本完成转换

注意事项

  • Contentful RichText有固定的节点类型规范,转换时需严格对应,避免无效节点
  • 复杂HTML(如嵌套标签、自定义属性)需要针对性处理节点的data字段

内容的提问来源于stack exchange,提问作者Eagleye

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.06 12:10:50