You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup在指定class的article区块内查找H1-H4标签

解决方法

你现在直接在整个soup对象上调用find_all,会匹配全页面所有符合规则的标签,只要先把查找范围限定在目标article区块内即可,操作如下:

写法1:先定位父容器再子查询(更易理解)

# 第一步:定位class为Article-p6ncbx-0 hxYamq的article主内容区块
target_article = soup.find("article", class_="Article-p6ncbx-0 hxYamq")

# 加判断避免未找到对应区块时报错
if target_article:
    # 仅在目标article范围内查找h1~h4标签
    for heading in target_article.find_all(["h1", "h2", "h3","h4"]):
        print(f"{heading.name} {heading.text.strip()}")

说明:class_是BeautifulSoup专门用于匹配HTMLclass属性的参数,加下划线是为了和Python的关键字class做区分。

写法2:CSS选择器写法(更简洁)

# 直接用CSS选择器规则匹配目标区块内的H标签
for heading in soup.select("article.Article-p6ncbx-0.hxYamq h1, article.Article-p6ncbx-0.hxYamq h2, article.Article-p6ncbx-0.hxYamq h3, article.Article-p6ncbx-0.hxYamq h4"):
    print(f"{heading.name} {heading.text.strip()}")

两种写法都可以实现仅匹配目标文章区块内的H标签,不会命中页头、页脚等外部区域的对应标签。

内容的提问来源于stack exchange,提问作者Daniel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.06 00:24:04