使用Python的BeautifulSoup无法获取h2标签值的求助
问题说明
尝试从larepublica.pe网站的「Economia」板块获取截图中显示的所有标题,但现有代码无法成功,疑似网页存在反爬限制。

现有代码
html = client.get("http://larepublica.pe/") soup = BeautifulSoup(html.text, 'html.parser') # Obtener la noticia de portada principal economyNews = "" for div in soup.findAll('h2', attrs={'class':'ItemSection_itemSection__title__PleA9'}): n = div.text economyNews += n+"\n" print(economyNews )
解决思路
1. 补全请求头,模拟浏览器访问
多数网站会通过请求头识别非浏览器请求,直接用client.get发送请求会被拦截。需添加完整的请求头参数:
import requests headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8', 'Accept-Language': 'es-ES,es;q=0.8' } html = requests.get("http://larepublica.pe/economia", headers=headers)
2. 处理动态渲染内容
如果目标标题是通过JavaScript动态加载的,静态HTML请求无法获取到内容,需用浏览器自动化工具模拟加载:
from selenium import webdriver from selenium.webdriver.common.by import By from bs4 import BeautifulSoup from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC driver = webdriver.Chrome() driver.get("http://larepublica.pe/economia") # 显式等待目标元素加载完成 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "ItemSection_itemSection__title__PleA9")) ) soup = BeautifulSoup(driver.page_source, 'html.parser') economyNews = "\n".join([h2.text for h2 in soup.find_all('h2', class_='ItemSection_itemSection__title__PleA9')]) print(economyNews) driver.quit()
3. 验证元素选择器有效性
网页结构可能已更新,原有的class名可能失效。打开浏览器开发者工具(F12),定位目标标题元素,确认最新的class、标签或父节点结构,调整选择器。
4. 使用会话保持Cookie
部分网站需要Cookie验证才能返回完整内容,用Session对象保持会话状态:
session = requests.Session() session.headers.update(headers) # 先发起一次请求获取网站Cookie session.get("http://larepublica.pe/") # 再请求Economia板块 html = session.get("http://larepublica.pe/economia")
5. 直接访问目标板块
原代码请求的是首页,直接访问Economia板块的URL(http://larepublica.pe/economia),能减少页面结构定位的复杂度,精准获取目标内容。
内容的提问来源于stack exchange,提问作者FreddicMatters
相关产品推荐
相关产品推荐

