Selenium网页爬取:点击加载更多后无法获取更新数据
我用Python结合Selenium爬取Burpple网站数据,点击“load more”按钮后,只能拿到初始数据,加载后的新数据获取不到,多次尝试都没解决。
我的代码
class RunChromeTests(): def test(self): chrome_options = Options() chrome_options.add_argument("--disable-notifications") chrome_options.add_argument("-incognito") chrome_options.add_argument("--disable-popup-blocking") chrome_options.add_argument("--ignore-certificate-errors") chrome_options.add_argument("--disable-javascript") # Download the chrome driver from https://chromedriver.chromium.org/downloads # and find the driver location in your computer chrome_path = r"path" driver = webdriver.Chrome(chrome_path, options=chrome_options) driver.maximize_window() driver.implicitly_wait(10) final = [] ### Enter your url to scrape but change the page number to {a} driver.get("https://www.burpple.com/search/sg?q=Newly+Opened&type=places") content = driver.page_source soup = BeautifulSoup(content) loadmore = driver.find_element_by_id("masonryViewMore-btn") j = 0 final1=[] try: while loadmore.is_displayed(): loadmore.click() time.sleep(2) lrec = soup.find_all("span",{"searchVenue-header-name-name headingMedium"}) #loadmore.is_displayed() newlist = lrec[j:] print(lrec) #print(newlist) for rec in newlist: name = rec.text #print(name) final1.append(name) print(final1) j = len(lrec)+1 #final1.append(name) time.sleep(5) #print(j) # except exceptions.StaleElementReferenceException: pass chromed = RunChromeTests() chromed.test()
当前输出
['Kotuwa', 'Smoochie Creamery', "Evan's Kitch", 'Plus Coffee Joint', '800° Woodfired Pizza (KINEX)', "Sarah's Loft", 'nicher (Springleaf)', 'First Story Cafe', 'Tucela Gelato', 'Ri Ri Cha', 'Unatoto', 'Royal Palm (Meat & Dine)']
期望输出
['Kotuwa', 'Smoochie Creamery', "Evan's Kitch", 'Plus Coffee Joint', '800° Woodfired Pizza (KINEX)', "Sarah's Loft", 'nicher (Springleaf)', 'First Story Cafe', 'Tucela Gelato', 'Ri Ri Cha', 'Unatoto', 'Royal Palm (Meat & Dine)', 'Flourish Bakehouse','Equate Coffee (Orchard Central)', 'TAG Espresso (Raffles City)', 'Arc-En-Ciel Patisserie', 'Enjoy Eating House & Bar (Stevens)','SAGE By Yasunori Doi (Orchard Plaza)','Hellu Coffee','Pestle & Mortar Society',... ]
核心问题点
- 禁用了JavaScript:代码里加了
--disable-javascript参数,而Burpple的Load More是通过JS动态加载内容的,禁用JS后点击按钮根本不会加载新数据。 - Soup未更新:只在页面初始化时获取了一次
page_source并创建soup,点击Load More后页面内容更新了,但soup还是旧的,自然拿不到新数据。 - 元素查找语法错误:
soup.find_all("span",{"searchVenue-header-name-name headingMedium"})写法错误,字典里没有指定class键,应该用class_参数或者{"class": "xxx"}。 - Load More元素失效:页面更新后,原来的
loadmore元素会变成陈旧元素(StaleElement),继续用原来的对象会报错,需要每次循环重新查找。
修改后的代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time class RunChromeTests(): def test(self): chrome_options = Options() chrome_options.add_argument("--disable-notifications") chrome_options.add_argument("-incognito") chrome_options.add_argument("--disable-popup-blocking") chrome_options.add_argument("--ignore-certificate-errors") # 移除禁用JS的参数 # chrome_options.add_argument("--disable-javascript") chrome_path = r"path" driver = webdriver.Chrome(chrome_path, options=chrome_options) driver.maximize_window() driver.implicitly_wait(10) final1 = [] driver.get("https://www.burpple.com/search/sg?q=Newly+Opened&type=places") while True: try: # 每次循环重新获取页面内容并初始化soup content = driver.page_source soup = BeautifulSoup(content, "html.parser") # 正确查找class对应的span元素 lrec = soup.find_all("span", class_="searchVenue-header-name-name headingMedium") # 提取新数据(避免重复添加) current_names = [rec.text.strip() for rec in lrec] # 只添加之前没出现过的名字 new_names = [name for name in current_names if name not in final1] final1.extend(new_names) print(f"当前已获取{len(final1)}条数据") print(final1) # 等待Load More按钮可点击,然后点击 loadmore = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.ID, "masonryViewMore-btn")) ) loadmore.click() # 等待新内容加载完成 time.sleep(3) except Exception as e: # 没有更多Load More按钮时退出循环 print("没有更多数据可加载,退出") break driver.quit() print("最终结果:") print(final1) chromed = RunChromeTests() chromed.test()
关键修改说明
- 移除了
--disable-javascript参数,确保动态加载功能正常。 - 每次循环都重新获取
page_source并创建新的soup,保证拿到最新页面内容。 - 修正了BeautifulSoup的元素查找方式,用
class_指定多类名。 - 使用
WebDriverWait等待按钮可点击,替代固定time.sleep,更可靠。 - 通过对比已有数据避免重复添加,确保最终列表没有重复项。
- 捕获所有异常,当按钮不可点击时退出循环,处理没有更多数据的情况。
内容的提问来源于stack exchange,提问作者Jingrui Lian

