You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取问题:指定class的div无法被BeautifulSoup捕获

问题:无法定位指定class的div元素

我尝试爬取带有如下属性的div元素:

<div data-v-28872a74="" class="col-lg-10 col-md-10 col-sm-12 col-12 offset-lg-1 offset-md-1 offset-sm-0 offset-0">

使用BeautifulSoup的find_all('div', class_ = 'col-lg-10 col-md-10 col-sm-12 col-12 offset-lg-1 offset-md-1 offset-sm-0 offset-0')调用后,返回空列表[]。

第一份测试代码(直接用requests请求)

import requests
from bs4 import BeautifulSoup as bs
url = 'https://remart.az/yasayis-kompleksi?cities=1&districts='

result = requests.get(url)
soup = bs(result.text, 'html.parser')
code= soup.find_all('div', class_ = 'col-lg-10 col-md-10  col-sm-12 col-12  offset-lg-1 offset-md-1 offset-sm-0 offset-0')
print(code)

第二份测试代码(Selenium获取列表URL,但详情页爬取仍失败)

先获取列表页的详情URL:

driver = webdriver.Chrome(r'C:\Program Files (x86)\chromedriver_win32\chromedriver.exe')
driver.get('https://remart.az/yasayis-kompleksi?cities=1&districts=')
time.sleep(3)

aze = driver.find_element(By.XPATH, '//*[@id="app"]/div[2]/div[1]/div[2]/div[6]/button')

for a in range(1,2):
    aze.click()
    time.sleep(1)

soup = bs(driver.page_source, "html.parser")
aezexx = soup.find_all('div', class_ = 'bitem')
for parent in aezexx:
    a_tag = parent.find("a")
    URRL = a_tag.attrs['href']
    print(URRL)

后续爬取详情页时,同样出现元素定位失败:

soup = bs(driver.page_source, "html.parser")
aezexx = soup.find_all('div', class_ = 'bitem')
for parent in aezexx:
    a_tag = parent.find("a")
    URRL = a_tag.attrs['href']
    result = requests.get(URRL)
    soup = bs(result.text, 'html.parser')
    are = soup.find_all("div", class_ = 'bottom-panel-descripton cut-text')
    for aes in are:
        azzzz = aes.find_all('p')
        print(azzzz) 

解决方案

1. 问题本质

  • 直接用requests请求只能拿到服务端返回的静态HTML骨架,目标元素是前端通过JavaScript动态渲染生成的,静态页面里根本没有这些元素,所以find_all找不到内容。
  • 详情页用requests请求也会踩同样的坑,因为详情页的核心内容同样是JS动态加载的。

2. 具体解决办法

办法一:全程用Selenium渲染页面

不管列表页还是详情页,都用Selenium加载页面(确保JS执行完毕)后再解析,这样能拿到完整的渲染后页面:

from selenium import webdriver
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup as bs
import time

# 初始化Chrome浏览器
driver = webdriver.Chrome(r'C:\Program Files (x86)\chromedriver_win32\chromedriver.exe')
# 打开列表页
driver.get('https://remart.az/yasayis-kompleksi?cities=1&districts=')
time.sleep(3)  # 等待页面加载完成

# 点击加载更多按钮(按需执行)
aze = driver.find_element(By.XPATH, '//*[@id="app"]/div[2]/div[1]/div[2]/div[6]/button')
for _ in range(1):  # 这里控制点击次数,原代码是1次
    aze.click()
    time.sleep(1)

# 解析列表页,提取所有详情页URL
soup = bs(driver.page_source, "html.parser")
detail_urls = [item.find("a")["href"] for item in soup.find_all('div', class_='bitem')]

# 遍历每个详情页,用Selenium加载后解析
for url in detail_urls:
    driver.get(url)
    time.sleep(2)  # 等待详情页加载
    soup = bs(driver.page_source, "html.parser")
    # 定位目标元素
    descriptions = soup.find_all("div", class_='bottom-panel-descripton cut-text')
    for desc in descriptions:
        paragraphs = desc.find_all('p')
        print(paragraphs)

# 关闭浏览器
driver.quit()

办法二:简化class匹配规则

不需要完全匹配所有class属性值,挑一个唯一且稳定的class片段来定位,比如:

# 只匹配关键class,避免空格或顺序问题
code = soup.find_all('div', class_='col-lg-10 offset-lg-1')

# 或者结合属性过滤,更精准
code = soup.find_all('div', {'data-v-28872a74': '', 'class': lambda cls: 'col-lg-10' in cls})

办法三:直接请求后端API(最优解)

打开浏览器开发者工具的「网络」面板,筛选XHR/fetch请求,找到页面加载数据的接口(比如类似/api/projects的请求),直接请求这些API获取JSON格式的数据,比解析HTML效率高得多,也避免了JS渲染的问题。


内容的提问来源于stack exchange,提问作者Rasim Dilbani

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 08:20:33