关于Google Maps爬虫获取经纬度与页面显示不一致的技术问询
如何获取Google Maps页面显示的准确经纬度数据?
我开发的网页爬虫从Google Maps抓取商家数据并保存至Excel文件,但Excel中的纬度(latitude)和经度(longitude)始终与Google Maps页面显示的实际值不符。请问如何获取与Google Maps页面完全一致的经纬度数据?
以下是我的代码:
def run_scraper(search_term: str, total_results: int = 100) -> list[dict]: business_list = BusinessList() with sync_playwright() as p: browser = p.chromium.launch(headless=True) page = browser.new_page() page.goto("https://www.google.com/maps", timeout=60000) print(f"Searching for: {search_term}") search_input = page.locator('//input[@id="searchboxinput"]') search_input.fill(search_term) page.keyboard.press("Enter") # Wait for initial results page.wait_for_selector('//a[contains(@href, "https://www.google.com/maps/place")]', timeout=20000) scrollable = page.locator('div[role="main"] div[aria-label]') scroll_pause_time = 1.5 prev_count = 0 max_attempts = 20 attempts = 0 # Scroll loop until we get at least `total_results` or exhaust attempts while True: page.mouse.wheel(0, 5000) time.sleep(scroll_pause_time) all_cards = page.locator('//div[contains(@class, "Nv2PK")]') count = all_cards.count() if count >= total_results or attempts >= max_attempts: break if count == prev_count: attempts += 1 else: attempts = 0 prev_count = count print(f"Total business cards loaded: {count}") listings = page.locator('//div[contains(@class, "Nv2PK")]').all()[:total_results] for i, listing in enumerate(listings, start=1): try: listing.scroll_into_view_if_needed() listing.click() time.sleep(4) business = Business() try: business.name = page.locator('//h1[contains(@class, "lfPIob")]').inner_text() except: business.name = "N/A" try: business.address = page.locator( '//button[@data-item-id="address"]//div[contains(@class, "fontBodyMedium")]' ).inner_text() except: business.address = "N/A" try: business.phone_number = page.locator( '//button[contains(@data-item-id, "phone:tel:")]//div[contains(@class, "fontBodyMedium")]' ).inner_text() except: business.phone_number = "N/A" try: business.website = page.locator('//a[@data-item-id="authority"]').get_attribute("href") or "N/A" except: business.website = "N/A" business.latitude, business.longitude = extract_coordinates_from_url(page.url) business_list.business_list.append(business) print(f"Scraped {i}: {business.name}, {business.address}, {business.phone_number}, {business.website}") except Exception as e: print(f"Error scraping listing {i}: {e}") continue browser.close()
问题根源
你当前通过extract_coordinates_from_url(page.url)从页面URL提取经纬度,但Google Maps页面URL中的坐标是近似值(用于简化URL或隐私保护),并非商家的精确坐标,这就是数据不符的核心原因。
两种可靠解决方法
方法1:从页面全局JS变量提取精确坐标
Google Maps会将精确坐标存储在页面的window.APP_INITIALIZATION_STATE全局变量中,通过Playwright执行JS代码即可获取:
# 替换原代码中提取坐标的那一行 coords = page.evaluate('''() => { const initState = window.APP_INITIALIZATION_STATE; if (initState && initState[3] && initState[3][6] && initState[3][6][0]) { return [initState[3][6][0][1], initState[3][6][0][2]]; // 纬度、经度 } return null; }''') if coords: business.latitude, business.longitude = coords else: business.latitude, business.longitude = "N/A", "N/A"
方法2:从DOM元素的属性提取
部分商家详情页的地图容器会在data-coordinates属性中存储精确坐标,直接读取该属性即可:
# 替换原代码中提取坐标的那一行 map_element = page.locator('div[data-coordinates]') if map_element.count() > 0: coord_str = map_element.get_attribute('data-coordinates') if coord_str: lat, lng = coord_str.split(',') business.latitude = lat.strip() business.longitude = lng.strip() else: business.latitude, business.longitude = "N/A", "N/A"
额外优化建议
- 把固定的
time.sleep(4)替换为Playwright的智能等待,比如page.wait_for_selector('h1.lfPIob', timeout=5000),既保证页面加载完成,又避免不必要的等待时间。 - Google Maps的页面结构和JS变量可能会随版本更新变化,若上述方法失效,可打开浏览器开发者工具(F12),在控制台搜索
latitude或longitude,定位最新的坐标存储位置。
内容的提问来源于stack exchange,提问作者mihir soni
相关产品推荐
相关产品推荐

