如何用BeautifulSoup提取含子标签的标签内容(不含父标签)
问题:使用BeautifulSoup提取标签内部内容(保留子标签但不含父标签)
需求是提取HTML标签的内容,不包含该标签本身,但保留内部所有子标签。示例代码如下:
html_text = "<TD WIDTH=60%>Portsmouth - Cherbourg<BR/>Portsmouth - Santander<BR/></TD>" soup = BeautifulSoup(html_text, 'html.parser') soup_list = soup.find_all("td") soup_object = soup_list[0] text = soup_object.getText() print(soup_object)
执行后输出带<td>标签的完整内容:
Portsmouth - Cherbourg
Portsmouth - Santander
但需要的输出是:
Portsmouth - Cherbourg
Portsmouth - Santander
使用soup_object.getText()会返回无<br/>标签的纯文本:
Portsmouth - CherbourgPortsmouth - Santander
soup_object.contents返回列表形式,也不符合需求:
['Portsmouth - Cherbourg',
, 'Portsmouth - Santander',
]
希望用BeautifulSoup的功能实现,不想用正则表达式。
解决方案
有两种直接用BeautifulSoup实现的方法:
方法1:使用decode_contents()方法
BeautifulSoup的Tag对象自带decode_contents()方法,可直接返回当前标签内部所有内容的HTML字符串,不包含标签本身。
示例代码:
html_text = "<TD WIDTH=60%>Portsmouth - Cherbourg<BR/>Portsmouth - Santander<BR/></TD>" soup = BeautifulSoup(html_text, 'html.parser') td_tag = soup.find("td") # 获取内部HTML内容 inner_html = td_tag.decode_contents() print(inner_html)
输出结果:
Portsmouth - Cherbourg
Portsmouth - Santander
方法2:遍历子节点并拼接
遍历目标标签的所有子节点,将每个子节点转为字符串后拼接,同样能得到不含父标签的内部HTML内容。
示例代码:
html_text = "<TD WIDTH=60%>Portsmouth - Cherbourg<BR/>Portsmouth - Santander<BR/></TD>" soup = BeautifulSoup(html_text, 'html.parser') td_tag = soup.find("td") # 拼接所有子节点的字符串形式 inner_html = ''.join(str(child) for child in td_tag.children) print(inner_html)
输出结果和方法1一致:
Portsmouth - Cherbourg
Portsmouth - Santander
内容的提问来源于stack exchange,提问作者Mangiafoco
相关产品推荐
相关产品推荐

