You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

markdownify.MarkdownConverter中convert_hN方法无法修改行为的问题咨询

markdownify.MarkdownConverter中convert_hN方法无法修改行为的问题咨询

你遇到的这个问题其实是两个小细节导致的,咱们来逐个解决:

首先,markdownify的默认实现里,h1和h2有独立的转换方法——也就是convert_h1和convert_h2,它们并不会调用通用的convert_hN方法。所以你只重写convert_hN的话,只有h3到h6的标题会走你的自定义逻辑,h1和h2完全不受影响,这就是为什么它们的输出里没有"BUG:"前缀。

其次,你在调用父类convert_hN的时候,错误地传了字符串"text"而不是变量text,这会导致父类方法拿不到标题的实际文本内容,就算逻辑对了也生成不了正确的结果。

这里给你两种修复方案,你可以根据需求选:

方案一:统一所有标题的转换逻辑

让h1和h2也调用我们重写的convert_hN,这样所有标题的处理逻辑保持一致:

from markdownify import MarkdownConverter

class CustomMarkdownConverter(MarkdownConverter):
    def convert_hN(self, n, el, text, parent_tags):
        # 调用父类方法获取原始Markdown格式
        title = super().convert_hN(n, el, text, parent_tags)
        return f"BUG: {title}"
    
    # 重写h1和h2的方法,让它们复用convert_hN的逻辑
    def convert_h1(self, el, text, parent_tags):
        return self.convert_hN(1, el, text, parent_tags)
    
    def convert_h2(self, el, text, parent_tags):
        return self.convert_hN(2, el, text, parent_tags)

def custom_markdownify(html):
    return CustomMarkdownConverter().convert(html)

html = """
<h1>Section 1</h1>
<h2>Section 2</h2>
<h3>Sub Section 3.1</h3>
"""
print(custom_markdownify(html))

运行后输出就会符合你的预期:

BUG: Section 1

BUG: Section 2

BUG: Sub Section 3.1

方案二:保留h1/h2的下划线样式,仅添加前缀

如果你想保持h1和h2原来的下划线分隔样式,只是加上前缀,可以单独重写这两个方法:

from markdownify import MarkdownConverter

class CustomMarkdownConverter(MarkdownConverter):
    def convert_h1(self, el, text, parent_tags):
        # 构造带前缀的标题文本
        prefixed_text = f"BUG: {text}"
        # 生成对应长度的下划线
        return f"{prefixed_text}\n" + "=" * len(prefixed_text)
    
    def convert_h2(self, el, text, parent_tags):
        prefixed_text = f"BUG: {text}"
        return f"{prefixed_text}\n" + "-" * len(prefixed_text)
    
    def convert_hN(self, n, el, text, parent_tags):
        # 对h3-h6,修改父类生成的#号格式标题
        original_title = super().convert_hN(n, el, text, parent_tags)
        return original_title.replace(text, f"BUG: {text}")

def custom_markdownify(html):
    return CustomMarkdownConverter().convert(html)

这个方案的输出会保留h1/h2的下划线样式,同时所有标题都加上"BUG:"前缀。

总结一下,核心问题就是默认的h1/h2不调用convert_hN,加上参数传递的笔误,导致你的自定义逻辑没生效。修正这两点就可以解决啦!

内容来源于stack exchange

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.08 14:28:05