谷歌地图餐厅Web Scraping程序返回空字符串数组,如何修复?
谷歌地图餐厅名称抓取返回空字符串问题
我正在开发一个Python项目,需要从谷歌地图抓取餐厅详情(名称、位置、价格、评分等)并存入SQL数据库。测试阶段仅尝试抓取餐厅名称存入record数组,但运行程序后返回多个空字符串组成的数组,怀疑问题出在获取名称的span标签定位上。
完整测试代码
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.common.keys import Keys from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver import ActionChains from selenium.webdriver.common.actions.wheel_input import ScrollOrigin from bs4 import BeautifulSoup import mysql.connector import time import requests service = Service(executable_path="chromedriver.exe") driver = webdriver.Chrome(service=service) wait = WebDriverWait(driver,10) count = 0 e_store = [] record = [] class Restaurants(): #stores the values from the google maps def __init__(self): self.name = None self.address = None self.rating = None self.website = None self.phone = None def open_browser(self): #To open the chrome web browser driver.maximize_window() driver.get('https://www.google.com/maps') try: but = wait.until(EC.element_to_be_clickable((By.XPATH, "//button[@aria-label='Accept all']"))) # make sure to make your input into a single tuple but.click() except Exception as e: print("Element not found", e) def search_term(self): add_term = driver.find_element(By.XPATH, "//input[@name='q']") ActionChains(driver)\ .send_keys_to_element(add_term, "Restaurants in my area")\ .key_down(Keys.RETURN)\ .perform() time.sleep(5) def get_all(self): global count #looking for restaurant profile terms = driver.find_elements(By.XPATH, "//a[@class='hfpxzc']") #This allows selenium to simulate user actions action = ActionChains(driver) # I think this is the code for scrolling while len(terms) < 1000: length = len(terms) #scroll away from staring profile scroll_origin = ScrollOrigin.from_element(terms[length-1]) action.scroll_from_origin(scroll_origin, 0, 1000).perform() #wait for 2 seconds time.sleep(2) terms = driver.find_elements(By.XPATH, "//a[@class='hfpxzc']") #Check if any new restraunt profile has appeared if len(terms) == length: count += 1 if count > 20: break else: count = 0 for i in range(len(terms)): scroll_origin = ScrollOrigin.from_element(terms[i]) action.scroll_from_origin(scroll_origin, 0,100).perform() action.move_to_element(terms[i]).perform() terms[i].click() time.sleep(2) source = driver.page_source soup = BeautifulSoup(source, 'html.parser') try: div_find = soup.find_all('div', class_="XltNde tTVLSc") for divs in div_find: name = divs.find('span', class_="a5H0ec") if name: record.append(name.text.strip()) print(record) except Exception as e: print("Failed:", e) continue # main program if __name__ == "__main__": get_details = Restaurants() get_details.open_browser() get_details.search_term() get_details.get_all()
运行结果
Cookie popup accepted [''] ['', ''] ['', '', ''] ['', '', '', ''] ['', '', '', '', ''] ['', '', '', '', '', '']
我猜测问题可能出在这段获取餐厅名称的代码上:
name = divs.find('span', class_="a5H0ec")
解决方案
问题根源
谷歌地图的页面元素类名是动态生成的,会频繁更新,你依赖的a5H0ec和XltNde tTVLSc这类类名已经失效,导致抓取到空内容。另外,直接用time.sleep()等待页面加载,可能还没等餐厅详情完全渲染就获取了页面源码,也会造成元素缺失。
修复步骤
- 替换为稳定的定位方式:放弃依赖动态类名,改用具有语义的XPath定位,比如通过元素的标签类型、文本特征或固定属性来定位。
- 优化等待逻辑:用Selenium的显式等待替代
time.sleep(),确保目标元素完全加载后再进行操作,避免因加载不完整导致的空值。 - 简化解析流程:直接用Selenium定位元素,不需要再转BeautifulSoup,减少中间环节的问题。
修复后的核心代码片段
修改get_all方法中的解析部分:
def get_all(self): global count #looking for restaurant profile terms = driver.find_elements(By.XPATH, "//a[@class='hfpxzc']") #This allows selenium to simulate user actions action = ActionChains(driver) # I think this is the code for scrolling while len(terms) < 1000: length = len(terms) #scroll away from staring profile scroll_origin = ScrollOrigin.from_element(terms[length-1]) action.scroll_from_origin(scroll_origin, 0, 1000).perform() #wait for 2 seconds time.sleep(2) terms = driver.find_elements(By.XPATH, "//a[@class='hfpxzc']") #Check if any new restraunt profile has appeared if len(terms) == length: count += 1 if count > 20: break else: count = 0 for i in range(len(terms)): scroll_origin = ScrollOrigin.from_element(terms[i]) action.scroll_from_origin(scroll_origin, 0,100).perform() action.move_to_element(terms[i]).perform() terms[i].click() # 改用显式等待,确保餐厅名称元素加载完成 try: # 定位餐厅名称的稳定XPath:匹配详情页的主标题 name_element = wait.until(EC.presence_of_element_located((By.XPATH, "//h1[contains(@class, 'DUwDvf')]"))) name = name_element.text.strip() if name: record.append(name) print(record) except Exception as e: print("Failed to get name:", e) continue
额外建议
- 避免使用全局变量
record,可以将其作为类的属性,代码结构更规范。 - 谷歌地图有反爬机制,频繁抓取可能触发验证码,建议增加随机等待时间,避免连续快速操作。
- 后续抓取其他字段(地址、评分等)时,同样使用显式等待+稳定定位:
- 地址:
//div[@data-item-id='address']//div[contains(@class, 'Io6YTe')] - 评分:
//div[@role='img']/@aria-label
- 地址:
内容的提问来源于stack exchange,提问作者itsALearningProcess
相关产品推荐
相关产品推荐

