如何用Selenium Python批量爬取Google Maps更多信息(无需点击条目)
修改后的Selenium代码(无需点击条目即可提取多字段)
以下代码基于你已有的get_coordinates_for_state函数,直接从Google Maps搜索结果列表中提取标题、地址、评分、评论数等信息,无需跳转至商家详情页:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.common.exceptions import NoSuchElementException import time # 你已实现的生成Google Maps URL的函数(保留你的原有逻辑即可) def get_coordinates_for_state(state): # 示例逻辑:替换为你实际生成URL的代码 return f"https://www.google.com/maps/search/圣诞树售卖点/@37.7749,-122.4194,12z/data=!3m1!4b1" def scrape_google_maps_christmas_trees(state): # 初始化Chrome驱动(可按需替换为Firefox等) options = webdriver.ChromeOptions() options.add_argument("--headless=new") # 无头模式,可选开启 options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options) driver.get(get_coordinates_for_state(state)) try: # 等待搜索结果容器加载完成 WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div[role='feed']")) ) # 滚动页面加载全部结果(Google Maps采用滚动加载机制) last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(3) new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 获取所有结果条目 results = driver.find_elements(By.CSS_SELECTOR, "div[role='feed'] > div:not([jsaction='mouseover:ignore;mousedown:ignore'])") scraped_data = [] for result in results: try: # 提取标题 title = result.find_element(By.CSS_SELECTOR, "div.fontHeadlineSmall").text # 提取地址 address = result.find_element(By.CSS_SELECTOR, "div.fontBodyMedium:nth-child(2)").text # 提取评分和评论数(处理无评分的情况) try: rating = result.find_element(By.CSS_SELECTOR, "span[role='img']").get_attribute("aria-label").split(" ")[0] review_count = result.find_element(By.CSS_SELECTOR, "div.fontBodyMedium:nth-child(3)").text.strip("()") except NoSuchElementException: rating = "无评分" review_count = "0" scraped_data.append({ "标题": title, "地址": address, "评分": rating, "评论数": review_count }) except NoSuchElementException: # 跳过格式异常的条目,避免程序崩溃 continue return scraped_data finally: # 确保驱动关闭 driver.quit() # 调用示例 if __name__ == "__main__": data = scrape_google_maps_christmas_trees("加利福尼亚州") for item in data: print(item)
核心说明
- 无需点击条目:直接定位搜索结果
feed容器下的每个条目,从条目内部提取所有字段,完全跳过详情页跳转步骤。 - 滚动加载处理:通过JS脚本滚动页面触发加载,确保获取全部搜索结果,而非仅默认显示的前几项。
- 异常兼容:针对部分商家无评分、评论的情况添加捕获逻辑,避免程序中断。
- 反爬优化:配置无头模式、禁用自动化检测、设置真实用户代理,降低被Google反爬拦截的概率。
注意事项
- Google Maps的页面元素选择器可能随版本更新变化,若某字段提取失败,需用浏览器开发者工具重新定位元素。
- 可根据需求扩展提取营业时间、联系方式等其他字段,只需在结果条目中找到对应元素并添加提取逻辑即可。
内容的提问来源于stack exchange,提问作者Yash
相关产品推荐
相关产品推荐

