You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Selenium网页爬取:点击加载更多后无法获取更新数据

问题:点击Load More后无法获取Burpple网站的新数据

我用Python结合Selenium爬取Burpple网站数据,点击“load more”按钮后,只能拿到初始数据,加载后的新数据获取不到,多次尝试都没解决。

我的代码

class RunChromeTests():
    def test(self):
        chrome_options = Options()
        chrome_options.add_argument("--disable-notifications")
        chrome_options.add_argument("-incognito")
        chrome_options.add_argument("--disable-popup-blocking")
        chrome_options.add_argument("--ignore-certificate-errors")
        chrome_options.add_argument("--disable-javascript")

    
        # Download the chrome driver from https://chromedriver.chromium.org/downloads 
        # and find the driver location in your computer
        chrome_path = r"path"
        driver = webdriver.Chrome(chrome_path, options=chrome_options)
        driver.maximize_window()
        driver.implicitly_wait(10)
        
        final = []
    
    
            
        ### Enter your url to scrape but change the page number to {a} 
        driver.get("https://www.burpple.com/search/sg?q=Newly+Opened&type=places")
        content = driver.page_source
        soup = BeautifulSoup(content)
            
        loadmore = driver.find_element_by_id("masonryViewMore-btn")
        j = 0
        final1=[]
       
        try:
            while loadmore.is_displayed():
                loadmore.click()
                time.sleep(2)
                lrec = soup.find_all("span",{"searchVenue-header-name-name headingMedium"})
                #loadmore.is_displayed()
                newlist = lrec[j:]
                print(lrec)
                #print(newlist)
                for rec in newlist:
                    name = rec.text
                    #print(name)
                    final1.append(name)
                print(final1)
                j = len(lrec)+1
                #final1.append(name)
                time.sleep(5)
                #print(j)
                #
        except exceptions.StaleElementReferenceException:
            pass
chromed = RunChromeTests()
chromed.test()

当前输出

['Kotuwa', 'Smoochie Creamery', "Evan's Kitch", 'Plus Coffee Joint', '800° Woodfired Pizza (KINEX)', "Sarah's Loft", 'nicher (Springleaf)', 'First Story Cafe', 'Tucela Gelato', 'Ri Ri Cha', 'Unatoto', 'Royal Palm (Meat & Dine)']

期望输出

['Kotuwa', 'Smoochie Creamery', "Evan's Kitch", 'Plus Coffee Joint', '800° Woodfired Pizza (KINEX)', "Sarah's Loft", 'nicher (Springleaf)', 'First Story Cafe', 'Tucela Gelato', 'Ri Ri Cha', 'Unatoto', 'Royal Palm (Meat & Dine)', 'Flourish Bakehouse','Equate Coffee (Orchard Central)', 'TAG Espresso (Raffles City)', 'Arc-En-Ciel Patisserie', 'Enjoy Eating House & Bar (Stevens)','SAGE By Yasunori Doi (Orchard Plaza)','Hellu Coffee','Pestle & Mortar Society',... ]


问题分析与解决方案

核心问题点

  1. 禁用了JavaScript:代码里加了--disable-javascript参数,而Burpple的Load More是通过JS动态加载内容的,禁用JS后点击按钮根本不会加载新数据。
  2. Soup未更新:只在页面初始化时获取了一次page_source并创建soup,点击Load More后页面内容更新了,但soup还是旧的,自然拿不到新数据。
  3. 元素查找语法错误:soup.find_all("span",{"searchVenue-header-name-name headingMedium"})写法错误,字典里没有指定class键,应该用class_参数或者{"class": "xxx"}。
  4. Load More元素失效:页面更新后,原来的loadmore元素会变成陈旧元素(StaleElement),继续用原来的对象会报错,需要每次循环重新查找。

修改后的代码

from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup
import time

class RunChromeTests():
    def test(self):
        chrome_options = Options()
        chrome_options.add_argument("--disable-notifications")
        chrome_options.add_argument("-incognito")
        chrome_options.add_argument("--disable-popup-blocking")
        chrome_options.add_argument("--ignore-certificate-errors")
        # 移除禁用JS的参数
        # chrome_options.add_argument("--disable-javascript")

        chrome_path = r"path"
        driver = webdriver.Chrome(chrome_path, options=chrome_options)
        driver.maximize_window()
        driver.implicitly_wait(10)
        
        final1 = []
    
        driver.get("https://www.burpple.com/search/sg?q=Newly+Opened&type=places")
        
        while True:
            try:
                # 每次循环重新获取页面内容并初始化soup
                content = driver.page_source
                soup = BeautifulSoup(content, "html.parser")
                # 正确查找class对应的span元素
                lrec = soup.find_all("span", class_="searchVenue-header-name-name headingMedium")
                
                # 提取新数据(避免重复添加)
                current_names = [rec.text.strip() for rec in lrec]
                # 只添加之前没出现过的名字
                new_names = [name for name in current_names if name not in final1]
                final1.extend(new_names)
                
                print(f"当前已获取{len(final1)}条数据")
                print(final1)
                
                # 等待Load More按钮可点击,然后点击
                loadmore = WebDriverWait(driver, 10).until(
                    EC.element_to_be_clickable((By.ID, "masonryViewMore-btn"))
                )
                loadmore.click()
                # 等待新内容加载完成
                time.sleep(3)
                
            except Exception as e:
                # 没有更多Load More按钮时退出循环
                print("没有更多数据可加载,退出")
                break
        
        driver.quit()
        print("最终结果:")
        print(final1)

chromed = RunChromeTests()
chromed.test()

关键修改说明

  • 移除了--disable-javascript参数,确保动态加载功能正常。
  • 每次循环都重新获取page_source并创建新的soup,保证拿到最新页面内容。
  • 修正了BeautifulSoup的元素查找方式,用class_指定多类名。
  • 使用WebDriverWait等待按钮可点击,替代固定time.sleep,更可靠。
  • 通过对比已有数据避免重复添加,确保最终列表没有重复项。
  • 捕获所有异常,当按钮不可点击时退出循环,处理没有更多数据的情况。

内容的提问来源于stack exchange,提问作者Jingrui Lian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.29 11:05:15