如何用BeautifulSoup 4.12.3获取包含table的<a>标签完整内容?
解决BeautifulSoup解析嵌套
<table>的<a>标签内容丢失问题 问题核心是HTML规范限制:<a>属于行内元素,标准中不允许嵌套<table>这类块级元素。BeautifulSoup默认使用的Python内置html.parser解析器会自动修正这种不符合规范的结构,把<table>从<a>标签内移出,导致你看到的<a>标签为空。
要保留原始嵌套结构,需换用对不规范HTML兼容性更好的解析器,比如lxml或html5lib:
步骤1:安装解析器
通过pip安装所需解析器(二选一即可):
pip install lxml # 或者 pip install html5lib
步骤2:修改代码指定解析器
初始化BeautifulSoup时明确指定解析器,就能保留<a>与<table>的嵌套关系:
import bs4 h = """ <a id="0"> <table> <thead> <tr> <th scope="col">Person</th> <th scope="col">Most interest in</th> <th scope="col">Age</th> </tr> </thead> <tbody> <tr> <th scope="row">Chris</th> <td>HTML tables</td> <td>22</td> </tr> </table> </a> """ # 使用lxml解析器 test = bs4.BeautifulSoup(h, 'lxml') # 或者用html5lib:test = bs4.BeautifulSoup(h, 'html5lib') # 获取包含完整table的a标签 a_tag = test.find('a') print(a_tag) # 单独提取关联的table内容 target_table = a_tag.find('table') print(target_table)
效果验证
修改后,test.find('a')会返回包含完整<table>的<a id="0">标签,你可以正常将id="0"与表格数据关联处理。
内容的提问来源于stack exchange,提问作者maggle
相关产品推荐
相关产品推荐

