You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

BeautifulSoup 4无法定位CSS与HTML标签问题求助

解决BeautifulSoup无法定位动态渲染标签的问题

核心原因

目标网站基于Angular框架开发,页面中的产品列表标签是通过JavaScript动态渲染生成的。requests.get()只能获取到初始的静态HTML源码,此时动态内容尚未加载,所以BeautifulSoup无法找到目标标签。

错误写法纠正

你之前尝试的几种定位方式存在语法错误,正确写法示例(但仅在静态页面存在该标签时有效):

  • CSS选择器:soup.select(".row.product-list")(类选择器需加.,多个类用.连接)
  • find_all方法:soup.find_all("div", class_="row product-list")(你的初始写法语法正确,但因为静态页面无此元素所以返回空)

有效解决方案

方案1:用Selenium模拟浏览器加载动态内容

Selenium可以模拟真实浏览器的行为,等待JavaScript渲染完成后再获取页面源码:

  1. 安装依赖:
pip install selenium
  1. 下载对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配),并确保驱动路径可被Python识别
  2. 示例代码:
from selenium import webdriver
from bs4 import BeautifulSoup
from selenium.webdriver.chrome.options import Options

url = 'https://qgold.com/pl/Jewelry-Rings-2%C2%B7Stone-Rings'
# 配置无头模式(可选,不弹出浏览器窗口)
options = Options()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)

driver.get(url)
# 等待页面加载完成,时间可根据实际情况调整
driver.implicitly_wait(5)

page_source = driver.page_source
soup = BeautifulSoup(page_source, "html.parser")
tovar = soup.find_all("div", class_="row product-list")
print(f"找到{len(tovar)}个目标标签")

driver.quit()

方案2:直接抓取数据接口(更高效)

动态渲染的内容通常来自后端API接口,可直接请求接口获取JSON格式数据,无需解析HTML:

  1. 打开浏览器开发者工具(F12),切换到「Network」标签,刷新页面
  2. 在「XHR/Fetch」分类下,查找返回产品列表的接口(通常URL包含product、api等关键词)
  3. 复制接口URL,添加必要请求头后直接请求:
import requests

api_url = "你找到的产品列表接口URL"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}
response = requests.get(api_url, headers=headers)
products = response.json()
# 按需处理产品数据
print(products)

注意事项

  • 使用Selenium时,需定期更新浏览器驱动以匹配浏览器版本
  • 抓取接口时,注意保留请求头中的关键参数(如User-Agent、Cookie),避免被网站拦截

内容的提问来源于stack exchange,提问作者Dima

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.20 07:13:25