You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Google Maps爬虫获取经纬度与页面显示不一致的技术问询

如何获取Google Maps页面显示的准确经纬度数据?

我开发的网页爬虫从Google Maps抓取商家数据并保存至Excel文件,但Excel中的纬度(latitude)和经度(longitude)始终与Google Maps页面显示的实际值不符。请问如何获取与Google Maps页面完全一致的经纬度数据?

以下是我的代码:

def run_scraper(search_term: str, total_results: int = 100) -> list[dict]:
    business_list = BusinessList()

    with sync_playwright() as p:
        browser = p.chromium.launch(headless=True)
        page = browser.new_page()
        page.goto("https://www.google.com/maps", timeout=60000)

        print(f"Searching for: {search_term}")
        search_input = page.locator('//input[@id="searchboxinput"]')
        search_input.fill(search_term)
        page.keyboard.press("Enter")

        # Wait for initial results
        page.wait_for_selector('//a[contains(@href, "https://www.google.com/maps/place")]', timeout=20000)

        scrollable = page.locator('div[role="main"] div[aria-label]')
        scroll_pause_time = 1.5  
        prev_count = 0
        max_attempts = 20
        attempts = 0

        # Scroll loop until we get at least `total_results` or exhaust attempts
        while True:
            page.mouse.wheel(0, 5000)
            time.sleep(scroll_pause_time)

            all_cards = page.locator('//div[contains(@class, "Nv2PK")]')
            count = all_cards.count()

            if count >= total_results or attempts >= max_attempts:
                break

            if count == prev_count:
                attempts += 1
            else:
                attempts = 0
                prev_count = count

        print(f"Total business cards loaded: {count}")

        listings = page.locator('//div[contains(@class, "Nv2PK")]').all()[:total_results]

        for i, listing in enumerate(listings, start=1):
            try:
                listing.scroll_into_view_if_needed()
                listing.click()
                time.sleep(4)

                business = Business()

                try:
                    business.name = page.locator('//h1[contains(@class, "lfPIob")]').inner_text()
                except:
                    business.name = "N/A"

                try:
                    business.address = page.locator(
                        '//button[@data-item-id="address"]//div[contains(@class, "fontBodyMedium")]'
                    ).inner_text()
                except:
                    business.address = "N/A"

                try:
                    business.phone_number = page.locator(
                        '//button[contains(@data-item-id, "phone:tel:")]//div[contains(@class, "fontBodyMedium")]'
                    ).inner_text()
                except:
                    business.phone_number = "N/A"

                try:
                    business.website = page.locator('//a[@data-item-id="authority"]').get_attribute("href") or "N/A"
                except:
                    business.website = "N/A"

                business.latitude, business.longitude = extract_coordinates_from_url(page.url)

                business_list.business_list.append(business)

                print(f"Scraped {i}: {business.name}, {business.address}, {business.phone_number}, {business.website}")

            except Exception as e:
                print(f"Error scraping listing {i}: {e}")
                continue

        browser.close()

问题根源

你当前通过extract_coordinates_from_url(page.url)从页面URL提取经纬度,但Google Maps页面URL中的坐标是近似值(用于简化URL或隐私保护),并非商家的精确坐标,这就是数据不符的核心原因。

两种可靠解决方法

方法1:从页面全局JS变量提取精确坐标

Google Maps会将精确坐标存储在页面的window.APP_INITIALIZATION_STATE全局变量中,通过Playwright执行JS代码即可获取:

# 替换原代码中提取坐标的那一行
coords = page.evaluate('''() => {
    const initState = window.APP_INITIALIZATION_STATE;
    if (initState && initState[3] && initState[3][6] && initState[3][6][0]) {
        return [initState[3][6][0][1], initState[3][6][0][2]]; // 纬度、经度
    }
    return null;
}''')
if coords:
    business.latitude, business.longitude = coords
else:
    business.latitude, business.longitude = "N/A", "N/A"

方法2:从DOM元素的属性提取

部分商家详情页的地图容器会在data-coordinates属性中存储精确坐标,直接读取该属性即可:

# 替换原代码中提取坐标的那一行
map_element = page.locator('div[data-coordinates]')
if map_element.count() > 0:
    coord_str = map_element.get_attribute('data-coordinates')
    if coord_str:
        lat, lng = coord_str.split(',')
        business.latitude = lat.strip()
        business.longitude = lng.strip()
else:
    business.latitude, business.longitude = "N/A", "N/A"

额外优化建议

  • 把固定的time.sleep(4)替换为Playwright的智能等待,比如page.wait_for_selector('h1.lfPIob', timeout=5000),既保证页面加载完成,又避免不必要的等待时间。
  • Google Maps的页面结构和JS变量可能会随版本更新变化,若上述方法失效,可打开浏览器开发者工具(F12),在控制台搜索latitude或longitude,定位最新的坐标存储位置。

内容的提问来源于stack exchange,提问作者mihir soni

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.12 23:30:03