You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python BeautifulSoup查找指定class元素仅返回首个结果问题求助

问题描述

爬取页面 https://www.britannica.com/place/Alabama-state 时,需要提取所有带 class="h1" 的h2标签内的文本。页面中至少存在2处符合规则的匹配元素:

  • 元素1:<section id="ref78299" data-level="1" data-has-spy="true"><h2 class="h1"><span id="ref613779"></span>Land</h2>,对应文本为 Land
  • 元素2:<section id="ref273744" data-level="1" data-has-spy="true"><h2 class="h1"><span id="ref613784"></span>People</h2>,对应文本为 People

使用如下代码提取时,仅输出第一个匹配结果Land,无法获取全部匹配项文本:

from bs4 import BeautifulSoup
import requests
htmlrequests=requests.get('https://www.britannica.com/place/Alabama-state')
htmlcontent=htmlrequests.content
soup=BeautifulSoup(htmlcontent,'html.parser')
for section in soup.find_all(class_='h1'):
    print(section.text)

当前实际输出仅为Land,预期输出为Land、People等所有匹配元素的文本。

问题原因
  • 默认requests请求没有携带合法的浏览器身份标识,触发网站反爬机制,返回的页面源码不完整,仅加载了首屏内容,后续章节的DOM节点未被返回
  • 选择器未限定标签类型,find_all(class_='h1')会匹配所有带h1类名的任意标签,可能匹配到无关节点干扰结果
解决方案

按以下两点调整代码即可拿到全部结果:

  1. 给请求添加浏览器User-Agent请求头,模拟正常用户访问,获取完整页面源码
  2. 精准限定匹配标签为h2,同时对提取的文本做空白字符清理,避免多余空行、空格影响输出

修正后可运行的完整代码:

from bs4 import BeautifulSoup
import requests

# 携带浏览器请求头,绕过基础反爬校验
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36"
}
resp = requests.get('https://www.britannica.com/place/Alabama-state', headers=headers)
soup = BeautifulSoup(resp.content, 'html.parser')

# 精准匹配class为h1的h2标签,提取文本时自动清理前后空白
for title in soup.find_all('h2', class_='h1'):
    print(title.get_text(strip=True))

运行代码后即可正常输出所有匹配的章节标题,包括Land、People等全部结果。

内容的提问来源于stack exchange,提问作者Faheem

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.27 12:18:14