You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用lxml etree打印完整HTML标签字符串而非仅标签名

解决方法

要输出完整的标签内容(含子节点和内部文本),使用lxml自带的etree.tostring()方法序列化节点即可,修改后的代码如下:

from bs4 import BeautifulSoup
from lxml import etree

doc = "<p><a></a><a></a>Printable Text"
soup = BeautifulSoup(doc, "lxml")
root = etree.fromstring(str(soup))

tree = etree.ElementTree(root)
for e in tree.iter():
    # 替换原有的e.tag打印逻辑
    print(etree.tostring(e, encoding='utf-8').decode('utf-8'))
    print("--------------")

代码说明

  • etree.tostring()会完整序列化当前节点的标签属性、所有子节点、内部文本,完全匹配你需要的输出格式
  • 显式指定encoding='utf-8'避免特殊字符乱码,再通过decode方法把返回的字节对象转为普通字符串,不会出现字节前缀b
    运行上述代码即可得到你预期的输出结果。

内容的提问来源于stack exchange,提问作者Andu Gundu Swami

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.03 09:42:01