You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

requests.get爬取异常:soup.find无法定位目标class求助

解决动态渲染页面的爬虫问题

问题原因

目标网站是单页应用(SPA),服务器返回的初始HTML仅包含页面框架,产品数据是通过JavaScript动态加载的。requests.get只能获取静态初始页面,无法捕获JS渲染后的内容,因此找不到df-card__title元素,且返回内容看起来不完整。

解决方案

方案1:使用Selenium模拟浏览器加载

Selenium可以模拟真实浏览器的渲染过程,获取完整的动态页面内容。

步骤:

  1. 安装Selenium:pip install selenium
  2. 下载对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配)

代码示例:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup

# 初始化Chrome驱动(确保驱动路径正确,或已配置到系统环境变量)
driver = webdriver.Chrome()
driver.get("https://farmacievigorito.it/#/dfclassic/query=cerotti&query_name=match_and")

# 等待目标元素加载完成,超时时间10秒
try:
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "df-card__title"))
    )
    # 获取渲染后的页面源码
    page_source = driver.page_source
    soup = BeautifulSoup(page_source, "html.parser")
    target_element = soup.find("div", class_="df-card__title")
    
    if target_element:
        print("产品名称:", target_element.get_text(strip=True))
    else:
        print("未找到目标元素")
finally:
    driver.quit()  # 关闭浏览器

方案2:直接调用API接口

通过浏览器开发者工具分析网站的API请求,直接获取产品数据(比解析HTML更高效)。

操作步骤:

  1. 打开浏览器开发者工具(F12),切换到「Network」标签
  2. 刷新页面,筛选「XHR/Fetch」类型的请求,找到加载产品列表的接口
  3. 复制接口URL、请求头和参数,用requests直接请求

代码示例(假设找到的API接口如下,需替换为实际接口):

import requests

api_url = "https://farmacievigorito.it/api/dfclassic/products?query=cerotti&query_name=match_and"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36",
    "Accept": "application/json"
}

response = requests.get(api_url, headers=headers)
if response.status_code == 200:
    product_data = response.json()
    # 根据实际JSON结构提取产品名称,示例结构:product_data["items"][0]["title"]
    if product_data.get("items"):
        print("产品名称:", product_data["items"][0]["title"])
    else:
        print("未找到产品数据")
else:
    print(f"API请求失败,状态码: {response.status_code}")

方案对比

  • Selenium:无需分析接口,适合复杂渲染场景,但运行速度慢、资源占用高
  • API调用:速度快、效率高,但需要手动分析接口参数和请求头,部分接口可能有反爬机制

内容的提问来源于stack exchange,提问作者Davide

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.25 19:12:36