You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python实现滚动加载页面爬取:获取至少50条车辆详情链接

爬取bama.ir滚动加载页面的车辆详情链接

问题背景

需要爬取https://bama.ir/car/samand-lx页面中所有车辆详情页的<a>标签链接(至少50条),但原代码仅能获取初始加载的3条链接——原因是该网站采用滚动加载动态渲染新内容,静态HTTP请求无法捕获后续加载的车辆数据。

原代码:

from bs4 import BeautifulSoup
import urllib.request
from http.cookiejar import CookieJar


baseUrl = 'https://bama.ir/car/'
brand = 'samand-lx'
url = f'{baseUrl}{brand}'

req = urllib.request.Request(url)
cj = CookieJar()
opener = urllib.request.build_opener(urllib.request.HTTPCookieProcessor(cj))
response = opener.open(req)

soup = BeautifulSoup(response, "html.parser")

links = soup.find_all('a', class_='bama-ad')
links = list(map(lambda x: 'https://bama.ir'+x['href'], links))
print(links)

原输出:

['https://bama.ir/car/detail-hbr06v3g-samand-lx-basic-1392', 'https://bama.ir/car/detail-pa4bsrlr-samand-lx-basic-1398', 'https://bama.ir/car/detail-dcuahpx9-samand-lx-ef7-1398']

解决方案:模拟浏览器滚动加载

使用selenium模拟浏览器滚动行为,触发动态加载逻辑,直到获取足够数量的链接。

前置准备

  1. 安装依赖:执行pip install selenium
  2. 下载对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配,可配置到系统环境变量或指定路径)

修改后代码

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import random

# 目标URL
url = 'https://bama.ir/car/samand-lx'

# 初始化Chrome浏览器(若驱动未在环境变量,需指定路径:webdriver.Chrome(executable_path="path/to/chromedriver"))
driver = webdriver.Chrome()
driver.get(url)

# 目标链接数量
target_count = 50
links = set()  # 用集合自动去重

while len(links) < target_count:
    # 提取当前页面所有符合条件的链接
    current_links = driver.find_elements(By.CSS_SELECTOR, 'a.bama-ad')
    for link in current_links:
        href = link.get_attribute('href')
        if href:
            links.add(href)
    
    # 模拟滚动到底部触发加载
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    
    # 随机等待1-3秒,避免触发反爬,同时给页面加载时间
    time.sleep(random.uniform(1, 3))
    
    # 等待新元素加载完成,超时则判定已无更多内容
    try:
        WebDriverWait(driver, 5).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, 'a.bama-ad:nth-last-child(1)'))
        )
    except:
        break

# 转换为列表格式
links_list = list(links)
print(f"共获取{len(links_list)}条链接:")
print(links_list[:5])  # 打印前5条示例

# 关闭浏览器
driver.quit()

核心细节说明

  • 用set存储链接,自动过滤重复的车辆详情页链接
  • 随机延迟滚动和等待,降低网站反爬机制的触发概率
  • 通过WebDriverWait等待新元素加载,确保滚动后动态内容已渲染完成
  • 若超时无法获取新元素,自动退出循环,避免无限等待

内容的提问来源于stack exchange,提问作者Sir-Sorg

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.23 06:52:39