markdownify.MarkdownConverter中convert_hN方法无法修改行为的问题咨询
markdownify.MarkdownConverter中convert_hN方法无法修改行为的问题咨询
你遇到的这个问题其实是两个小细节导致的,咱们来逐个解决:
首先,markdownify的默认实现里,h1和h2有独立的转换方法——也就是convert_h1和convert_h2,它们并不会调用通用的convert_hN方法。所以你只重写convert_hN的话,只有h3到h6的标题会走你的自定义逻辑,h1和h2完全不受影响,这就是为什么它们的输出里没有"BUG:"前缀。
其次,你在调用父类convert_hN的时候,错误地传了字符串"text"而不是变量text,这会导致父类方法拿不到标题的实际文本内容,就算逻辑对了也生成不了正确的结果。
这里给你两种修复方案,你可以根据需求选:
方案一:统一所有标题的转换逻辑
让h1和h2也调用我们重写的convert_hN,这样所有标题的处理逻辑保持一致:
from markdownify import MarkdownConverter class CustomMarkdownConverter(MarkdownConverter): def convert_hN(self, n, el, text, parent_tags): # 调用父类方法获取原始Markdown格式 title = super().convert_hN(n, el, text, parent_tags) return f"BUG: {title}" # 重写h1和h2的方法,让它们复用convert_hN的逻辑 def convert_h1(self, el, text, parent_tags): return self.convert_hN(1, el, text, parent_tags) def convert_h2(self, el, text, parent_tags): return self.convert_hN(2, el, text, parent_tags) def custom_markdownify(html): return CustomMarkdownConverter().convert(html) html = """ <h1>Section 1</h1> <h2>Section 2</h2> <h3>Sub Section 3.1</h3> """ print(custom_markdownify(html))
运行后输出就会符合你的预期:
BUG: Section 1BUG: Section 2
BUG: Sub Section 3.1
方案二:保留h1/h2的下划线样式,仅添加前缀
如果你想保持h1和h2原来的下划线分隔样式,只是加上前缀,可以单独重写这两个方法:
from markdownify import MarkdownConverter class CustomMarkdownConverter(MarkdownConverter): def convert_h1(self, el, text, parent_tags): # 构造带前缀的标题文本 prefixed_text = f"BUG: {text}" # 生成对应长度的下划线 return f"{prefixed_text}\n" + "=" * len(prefixed_text) def convert_h2(self, el, text, parent_tags): prefixed_text = f"BUG: {text}" return f"{prefixed_text}\n" + "-" * len(prefixed_text) def convert_hN(self, n, el, text, parent_tags): # 对h3-h6,修改父类生成的#号格式标题 original_title = super().convert_hN(n, el, text, parent_tags) return original_title.replace(text, f"BUG: {text}") def custom_markdownify(html): return CustomMarkdownConverter().convert(html)
这个方案的输出会保留h1/h2的下划线样式,同时所有标题都加上"BUG:"前缀。
总结一下,核心问题就是默认的h1/h2不调用convert_hN,加上参数传递的笔误,导致你的自定义逻辑没生效。修正这两点就可以解决啦!
内容来源于stack exchange
相关产品推荐
相关产品推荐

