You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:使用Selenium爬取Google Events仅获取20条数据的问题

问题

我正尝试爬取Google上特定时间段内的所有活动,目前代码虽能运行,但仅能获取50余条数据中的20条。尽管已使用滚动查看命令,仍无法抓取全部页面元素,请问问题出在哪里?

以下是我的代码:

from selenium import webdriver
from time import sleep
from selenium.common.exceptions import NoSuchElementException

chrome_driver_path = "/Development/chromedriver"
driver = webdriver.Chrome(executable_path=chrome_driver_path)

website = "https://www.google.com/search?q=charlotte+nc&ei=6ObaYqueDfWrqtsPsP2MuAs&uact=5&oq=events&gs_lcp=Cgdnd3Mtd2l6EAMyBAgAEEMyCggAELEDEIMBEEMyCggAELEDEMkDEEMyBQgAEJIDMgUIABCSAzILCAAQgAQQsQMQgwEyBAgAEEMyBQgAEIAEMgcIABCxAxBDMgoIABCxAxCDARBDOgoILhDHARDRAxBDOhAILhCxAxCDARDHARDRAxBDOhEILhCABBCxAxCDARDHARDRAzoICC4QgAQQsQNKBAhBGABKBAhGGABQAFjGBWCMB2gAcAF4AIABkwGIAYUGkgEDMS41mAEAoAEBwAEB&sclient=gws-wiz&ibp=htl;events&rciv=evn&sa=X&ved=2ahUKEwjV8dWAi435AhWenWoFHU5-BxAQ8eoFKAJ6BAgDEA8#fpstate=tldetail&htichips=date:week&htischips=&htivrt=events&htidocid=L2F1dGhvcml0eS9ob3Jpem9uL2NsdXN0ZXJlZF9ldmVudC8yMDIyLTA3LTI0fDE3NTc3NDU4NDA4OTA0MDcxMDk2"
driver.get(website)
sleep(2)


driver.find_element_by_xpath("//div[@jsname='SJ2nIb'][4]").click()
sleep(1)

for x in driver.find_elements_by_xpath("//div[@class='YOGjf' and @role='heading']"):

    try:
        x.click()
    except:
        pass
    sleep(2)
    try:
        name = driver.find_element_by_xpath("//div[@id='Hc1J4e']//div[@class='dEuIWb']").text
    except NoSuchElementException:
        name = driver.find_element_by_xpath("//div[@id='Hc1J4e']//div[@class='dEuIWb Nre7t']").text

    try:
        times = driver.find_element_by_xpath("//div[@id='Hc1J4e']//div[@class='Gkoz3']").text
    except NoSuchElementException:
        times = ""
        print(name,"no times")
    try:
        address = driver.find_element_by_xpath("//div[@id='Hc1J4e']//span[@class='U6txu']").text
    except NoSuchElementException:
        address = ""
        print(name,"no addy")
    try:
        description = driver.find_element_by_xpath("//div[@id='Hc1J4e']//span[@class='PVlUWc']").text
    except NoSuchElementException:
        description = ""
        print(name,"no description")
    try:
        tickets = driver.find_element_by_xpath("//div[@id='Hc1J4e']//a[@class='uaYYHd vQQnKe uMdbSb']").get_attribute("href")
    except:
        tickets = ""
        print(name,"no tickets")
    driver.execute_script('arguments[0].scrollIntoView();', x)
    print(name, times, address, description, tickets)

driver.quit()
解决分析

核心问题

  1. 初始元素列表仅包含已渲染的前20条:你在循环前执行driver.find_elements_by_xpath时,页面只加载了前20条活动,后续活动需要滚动到底部才会动态渲染。一旦获取了元素列表,后续新加载的元素不会自动加入列表,导致循环只处理最初的20条。
  2. 滚动逻辑未触发加载:你是在处理单个元素后滚动到该元素位置,这种滚动幅度不足以触发Google活动页面的加载更多逻辑,必须滚动到页面底部才会触发新内容加载。

修复方案

步骤1:先滚动加载所有活动

在获取元素列表前,循环滚动到页面底部,通过对比滚动前后的页面高度判断是否加载完成。

步骤2:重新获取完整元素列表

等所有活动加载完毕后,再获取全部活动标题元素,之后遍历处理。

修复后的示例代码

from selenium import webdriver
from time import sleep
from selenium.common.exceptions import NoSuchElementException

chrome_driver_path = "/Development/chromedriver"
driver = webdriver.Chrome(executable_path=chrome_driver_path)

website = "https://www.google.com/search?q=charlotte+nc&ei=6ObaYqueDfWrqtsPsP2MuAs&uact=5&oq=events&gs_lcp=Cgdnd3Mtd2l6EAMyBAgAEEMyCggAELEDEIMBEEMyCggAELEDEMkDEEMyBQgAEJIDMgUIABCSAzILCAAQgAQQsQMQgwEyBAgAEEMyBQgAEIAEMgcIABCxAxBDMgoIABCxAxCDARBDOgoILhDHARDRAxBDOhAILhCxAxCDARDHARDRAxBDOhEILhCABBCxAxCDARDHARDRAzoICC4QgAQQsQNKBAhBGABKBAhGGABQAFjGBWCMB2gAcAF4AIABkwGIAYUGkgEDMS41mAEAoAEBwAEB&sclient=gws-wiz&ibp=htl;events&rciv=evn&sa=X&ved=2ahUKEwjV8dWAi435AhWenWoFHU5-BxAQ8eoFKAJ6BAgDEA8#fpstate=tldetail&htichips=date:week&htischips=&htivrt=events&htidocid=L2F1dGhvcml0eS9ob3Jpem9uL2NsdXN0ZXJlZF9ldmVudC8yMDIyLTA3LTI0fDE3NTc3NDU4NDA4OTA0MDcxMDk2"
driver.get(website)
sleep(2)

driver.find_element_by_xpath("//div[@jsname='SJ2nIb'][4]").click()
sleep(1)

# 先滚动加载所有活动
last_height = driver.execute_script("return document.body.scrollHeight")
while True:
    # 滚动到页面底部
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    sleep(3)  # 等待加载时间,可根据网络情况调整
    new_height = driver.execute_script("return document.body.scrollHeight")
    # 高度不再变化则停止滚动
    if new_height == last_height:
        break
    last_height = new_height

# 获取全部活动标题元素
event_titles = driver.find_elements_by_xpath("//div[@class='YOGjf' and @role='heading']")

for x in event_titles:
    try:
        x.click()
    except:
        pass
    sleep(2)
    try:
        name = driver.find_element_by_xpath("//div[@id='Hc1J4e']//div[@class='dEuIWb']").text
    except NoSuchElementException:
        name = driver.find_element_by_xpath("//div[@id='Hc1J4e']//div[@class='dEuIWb Nre7t']").text

    try:
        times = driver.find_element_by_xpath("//div[@id='Hc1J4e']//div[@class='Gkoz3']").text
    except NoSuchElementException:
        times = ""
        print(name,"no times")
    try:
        address = driver.find_element_by_xpath("//div[@id='Hc1J4e']//span[@class='U6txu']").text
    except NoSuchElementException:
        address = ""
        print(name,"no addy")
    try:
        description = driver.find_element_by_xpath("//div[@id='Hc1J4e']//span[@class='PVlUWc']").text
    except NoSuchElementException:
        description = ""
        print(name,"no description")
    try:
        tickets = driver.find_element_by_xpath("//div[@id='Hc1J4e']//a[@class='uaYYHd vQQnKe uMdbSb']").get_attribute("href")
    except:
        tickets = ""
        print(name,"no tickets")
    print(name, times, address, description, tickets)

driver.quit()

额外优化建议

  • 替换sleep()为显式等待:使用WebDriverWait配合expected_conditions等待元素加载,比固定等待时间更可靠,避免因网络延迟导致的元素未加载问题。
  • 处理元素过期问题:滚动加载后,之前的元素可能变成过期状态,因此必须在加载完成后重新获取元素列表,这也是修复代码中先滚动再获取列表的原因。

内容的提问来源于stack exchange,提问作者Aggregate Score

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 20:06:33