如何使用lxml etree打印完整HTML标签字符串而非仅标签名
解决方法
要输出完整的标签内容(含子节点和内部文本),使用lxml自带的etree.tostring()方法序列化节点即可,修改后的代码如下:
from bs4 import BeautifulSoup from lxml import etree doc = "<p><a></a><a></a>Printable Text" soup = BeautifulSoup(doc, "lxml") root = etree.fromstring(str(soup)) tree = etree.ElementTree(root) for e in tree.iter(): # 替换原有的e.tag打印逻辑 print(etree.tostring(e, encoding='utf-8').decode('utf-8')) print("--------------")
代码说明
etree.tostring()会完整序列化当前节点的标签属性、所有子节点、内部文本,完全匹配你需要的输出格式- 显式指定
encoding='utf-8'避免特殊字符乱码,再通过decode方法把返回的字节对象转为普通字符串,不会出现字节前缀b
运行上述代码即可得到你预期的输出结果。
内容的提问来源于stack exchange,提问作者Andu Gundu Swami
相关产品推荐
相关产品推荐

