Python BeautifulSoup查找指定class元素仅返回首个结果问题求助
问题描述
爬取页面 https://www.britannica.com/place/Alabama-state 时,需要提取所有带 class="h1" 的h2标签内的文本。页面中至少存在2处符合规则的匹配元素:
- 元素1:
<section id="ref78299" data-level="1" data-has-spy="true"><h2 class="h1"><span id="ref613779"></span>Land</h2>,对应文本为 Land - 元素2:
<section id="ref273744" data-level="1" data-has-spy="true"><h2 class="h1"><span id="ref613784"></span>People</h2>,对应文本为 People
使用如下代码提取时,仅输出第一个匹配结果Land,无法获取全部匹配项文本:
from bs4 import BeautifulSoup import requests htmlrequests=requests.get('https://www.britannica.com/place/Alabama-state') htmlcontent=htmlrequests.content soup=BeautifulSoup(htmlcontent,'html.parser') for section in soup.find_all(class_='h1'): print(section.text)
当前实际输出仅为Land,预期输出为Land、People等所有匹配元素的文本。
问题原因
- 默认
requests请求没有携带合法的浏览器身份标识,触发网站反爬机制,返回的页面源码不完整,仅加载了首屏内容,后续章节的DOM节点未被返回 - 选择器未限定标签类型,
find_all(class_='h1')会匹配所有带h1类名的任意标签,可能匹配到无关节点干扰结果
解决方案
按以下两点调整代码即可拿到全部结果:
- 给请求添加浏览器
User-Agent请求头,模拟正常用户访问,获取完整页面源码 - 精准限定匹配标签为
h2,同时对提取的文本做空白字符清理,避免多余空行、空格影响输出
修正后可运行的完整代码:
from bs4 import BeautifulSoup import requests # 携带浏览器请求头,绕过基础反爬校验 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/125.0.0.0 Safari/537.36" } resp = requests.get('https://www.britannica.com/place/Alabama-state', headers=headers) soup = BeautifulSoup(resp.content, 'html.parser') # 精准匹配class为h1的h2标签,提取文本时自动清理前后空白 for title in soup.find_all('h2', class_='h1'): print(title.get_text(strip=True))
运行代码后即可正常输出所有匹配的章节标题,包括Land、People等全部结果。
内容的提问来源于stack exchange,提问作者Faheem
相关产品推荐
相关产品推荐

