You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Selenium Python批量爬取Google Maps更多信息(无需点击条目)

修改后的Selenium代码(无需点击条目即可提取多字段)

以下代码基于你已有的get_coordinates_for_state函数,直接从Google Maps搜索结果列表中提取标题、地址、评分、评论数等信息,无需跳转至商家详情页:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import NoSuchElementException
import time

# 你已实现的生成Google Maps URL的函数(保留你的原有逻辑即可)
def get_coordinates_for_state(state):
    # 示例逻辑:替换为你实际生成URL的代码
    return f"https://www.google.com/maps/search/圣诞树售卖点/@37.7749,-122.4194,12z/data=!3m1!4b1"

def scrape_google_maps_christmas_trees(state):
    # 初始化Chrome驱动(可按需替换为Firefox等)
    options = webdriver.ChromeOptions()
    options.add_argument("--headless=new")  # 无头模式,可选开启
    options.add_argument("--disable-blink-features=AutomationControlled")
    options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
    
    driver = webdriver.Chrome(options=options)
    driver.get(get_coordinates_for_state(state))
    
    try:
        # 等待搜索结果容器加载完成
        WebDriverWait(driver, 20).until(
            EC.presence_of_element_located((By.CSS_SELECTOR, "div[role='feed']"))
        )
        
        # 滚动页面加载全部结果(Google Maps采用滚动加载机制)
        last_height = driver.execute_script("return document.body.scrollHeight")
        while True:
            driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
            time.sleep(3)
            new_height = driver.execute_script("return document.body.scrollHeight")
            if new_height == last_height:
                break
            last_height = new_height
        
        # 获取所有结果条目
        results = driver.find_elements(By.CSS_SELECTOR, "div[role='feed'] > div:not([jsaction='mouseover:ignore;mousedown:ignore'])")
        scraped_data = []
        
        for result in results:
            try:
                # 提取标题
                title = result.find_element(By.CSS_SELECTOR, "div.fontHeadlineSmall").text
                
                # 提取地址
                address = result.find_element(By.CSS_SELECTOR, "div.fontBodyMedium:nth-child(2)").text
                
                # 提取评分和评论数(处理无评分的情况)
                try:
                    rating = result.find_element(By.CSS_SELECTOR, "span[role='img']").get_attribute("aria-label").split(" ")[0]
                    review_count = result.find_element(By.CSS_SELECTOR, "div.fontBodyMedium:nth-child(3)").text.strip("()")
                except NoSuchElementException:
                    rating = "无评分"
                    review_count = "0"
                
                scraped_data.append({
                    "标题": title,
                    "地址": address,
                    "评分": rating,
                    "评论数": review_count
                })
            except NoSuchElementException:
                # 跳过格式异常的条目,避免程序崩溃
                continue
        
        return scraped_data
    
    finally:
        # 确保驱动关闭
        driver.quit()

# 调用示例
if __name__ == "__main__":
    data = scrape_google_maps_christmas_trees("加利福尼亚州")
    for item in data:
        print(item)

核心说明

  • 无需点击条目:直接定位搜索结果feed容器下的每个条目,从条目内部提取所有字段,完全跳过详情页跳转步骤。
  • 滚动加载处理:通过JS脚本滚动页面触发加载,确保获取全部搜索结果,而非仅默认显示的前几项。
  • 异常兼容:针对部分商家无评分、评论的情况添加捕获逻辑,避免程序中断。
  • 反爬优化:配置无头模式、禁用自动化检测、设置真实用户代理,降低被Google反爬拦截的概率。

注意事项

  • Google Maps的页面元素选择器可能随版本更新变化,若某字段提取失败,需用浏览器开发者工具重新定位元素。
  • 可根据需求扩展提取营业时间、联系方式等其他字段,只需在结果条目中找到对应元素并添加提取逻辑即可。

内容的提问来源于stack exchange,提问作者Yash

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 22:05:18