新手求助:如何用BeautifulSoup解析h2及下属h3并提取指定h2文本
解决思路与代码示例
嘿,刚接触lxml和BeautifulSoup完全不用客气~新手入门都会有这样的小困惑,我来一步步帮你搞定!
一、获取指定h2标签的文本
你已经通过soup.find_all("h2")拿到了所有h2标签,接下来只需要遍历这些标签,判断里面<strong>的文本是否是你要找的目标即可:
from bs4 import BeautifulSoup # 假设你已经把CSV中的HTML内容解析成了soup对象 for h2_tag in soup.find_all("h2"): # 提取h2下strong标签的文本,strip()去除多余空格和换行 heading = h2_tag.find("strong").get_text(strip=True) # 判断是否是目标标题 if heading in ["Powers and Abilities", "Weapons and Equipment"]: print(f"找到目标标题:{heading}") # 如果需要获取该标题下的具体内容,还可以用h2_tag.next_sibling等方法遍历后续节点
二、解析h2标签及其下属的h3标签
通常这类文档结构中,h3是h2的同级子章节(紧跟在h2之后,直到下一个h2出现),我们可以通过遍历h2的兄弟节点来收集对应的h3:
h2_list = soup.find_all("h2") for h2 in h2_list: h2_title = h2.find("strong").get_text(strip=True) print(f"\n=== {h2_title} ===") # 从h2的下一个节点开始遍历 current_node = h2.next_sibling while current_node: # 跳过空白文本节点(HTML中的换行、空格等) if isinstance(current_node, str) and current_node.strip() == "": current_node = current_node.next_sibling continue # 遇到下一个h2就停止当前h2的遍历 if current_node.name == "h2": break # 如果是h3标签,提取文本 if current_node.name == "h3": h3_title = current_node.get_text(strip=True) print(f" - {h3_title}") # 移动到下一个节点 current_node = current_node.next_sibling
小提醒
因为你是从CSV文件中读取数据,记得先通过csv模块把对应单元格里的HTML字符串提取出来,再用BeautifulSoup(html_str, "lxml")创建解析对象哦~
内容的提问来源于stack exchange,提问作者Lunalight
相关产品推荐
相关产品推荐

