You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Google Map评论评分与评论数爬取异常:仅获取前4条数据

问题描述

使用Playwright爬取Google地图商家信息,计划爬取8条数据,但仅能提取前4条的评论评分(reviews_average)和评论数(reviews_count),剩余4条无法获取该类数据。

爬取代码

for index in range(min(total, len(listings))):
            try:
                listings[index].click()
                page.wait_for_timeout(8000)
            except playwright._impl._api_types.Error as e:
                print(f"Error clicking on listing {index + 1}: {e}")
                continue

            name_xpath = f'(//div[contains(@class, "qBF1Pd fontHeadlineSmall ")])[{index + 1}]'
            address_xpath = '//button[@data-item-id="address"]//div[contains(@class, "fontBodyMedium")]'
            website_xpath = '//a[@data-item-id="authority"]//div[contains(@class, "fontBodyMedium")]'
            phone_number_xpath = '//button[contains(@data-item-id, "phone:tel:")]//div[contains(@class, "fontBodyMedium")]'
            category_xpath = '//*[@id="QA0Szd"]/div/div/div[1]/div[3]/div/div[1]/div/div/div[2]/div[2]/div/div[1]/div[2]/div/div[2]/span/span/button'
            reviews_span_xpath = f'//div[{index + 1}]//span[@role="img"]'
            
           
            business = Business()

            if page.locator(name_xpath).count() > 0:
                business.name = page.locator(name_xpath).inner_text()
            else:
                business.name = "N/A"

            if page.locator(address_xpath).count() > 0:
                business.address = page.locator(address_xpath).inner_text()
            else:
                business.address = "N/A"

            if page.locator(website_xpath).count() > 0:
                business.website = page.locator(website_xpath).inner_text()
            else:
                business.website = "N/A"

            if page.locator(phone_number_xpath).count() > 0:
                business.phone_number = page.locator(phone_number_xpath).inner_text()
            else:
                business.phone_number = "N/A"

            if page.locator(category_xpath).count() > 0:
                business.category = page.locator(category_xpath).inner_text()
            else:
                business.category = "N/A"

            page.mouse.wheel(0, 20000)
            page.wait_for_timeout(1000)
            reviews_elements = listings[index].locator(reviews_span_xpath)  # Move this line inside the loop
            if reviews_elements.count() > 0:
                reviews_label = reviews_elements.nth(0).get_attribute("aria-label")
                print(f"Reviews Label: {reviews_label}")

                # Process the reviews label to extract count and average
                # Example processing using regular expressions
                import re
                match = re.match(r'([\d.]+) stars ([\d,]+) Reviews', reviews_label)

                if match:
                    reviews_average = float(match.group(1))
                    reviews_count = int(re.sub(',', '', match.group(2)))
                else:
                    reviews_average = None
                    reviews_count = None

                # Update the Business object with reviews information
                business.reviews_average = reviews_average
                business.reviews_count = reviews_count
            else:
                business.reviews_average = None
                business.reviews_count = None

输出结果

Total Scraped: 8
Reviews Label: 4.6 stars 117 Reviews
Reviews Label: 5.0 stars 1 Reviews
Reviews Label: 4.6 stars 1,308 Reviews
Reviews Label: 4.9 stars 1,021 Reviews
问题分析与解决办法

核心问题

  1. 全局XPATH定位失效:reviews_span_xpath用页面全局div索引定位,滚动页面后列表项DOM结构变化,索引不再对应目标商家。
  2. 滚动时机错误:点击商家后滚动整个页面,可能导致当前列表项DOM被移出视图或结构变更,无法定位评论元素。
  3. 未利用列表上下文:未基于listings[index]的上下文定位评论元素,全局XPATH易受页面动态结构影响。

修复步骤

  1. 改用相对路径定位评论元素:基于当前列表项的上下文查找,避免全局索引干扰:
    # 替换原reviews_span_xpath定义
    reviews_span_xpath = './/span[@role="img"]'  # 相对路径,仅在当前listings[index]范围内查找
    
  2. 优化等待机制:替换固定超时等待,改用Playwright的元素加载等待,确保页面就绪:
    # 替换page.wait_for_timeout(8000)
    listings[index].click()
    page.wait_for_selector('//button[@data-item-id="address"]', timeout=10000)
    
  3. 移除不必要的全局滚动:若评论在详情面板内,滚动详情面板而非整个页面,避免干扰列表DOM:
    # 替换page.mouse.wheel(0, 20000)
    detail_panel = page.locator('//div[@id="QA0Szd"]//div[contains(@class, "m6QErb")]')
    detail_panel.scroll_into_view_if_needed()
    
  4. 适配正则匹配格式:兼容不同地区的评论文本表述,避免格式不匹配导致提取失败:
    # 替换原正则匹配
    match = re.match(r'([\d.]+) stars? ([\d,]+) (?:Reviews?|评价)', reviews_label)
    

修复后的核心代码片段

for index in range(min(total, len(listings))):
    try:
        listings[index].click()
        # 等待详情面板加载完成
        page.wait_for_selector('//button[@data-item-id="address"]', timeout=10000)
    except playwright._impl._api_types.Error as e:
        print(f"点击第{index + 1}个商家出错: {e}")
        continue

    # 其他元素定位逻辑保持不变...

    # 调整评论元素定位为相对路径
    reviews_span_xpath = './/span[@role="img"]'
    reviews_elements = listings[index].locator(reviews_span_xpath)
    
    if reviews_elements.count() > 0:
        reviews_label = reviews_elements.nth(0).get_attribute("aria-label")
        print(f"评论标签: {reviews_label}")

        import re
        # 适配多语言评论文本格式
        match = re.match(r'([\d.]+) stars? ([\d,]+) (?:Reviews?|评价)', reviews_label)

        if match:
            reviews_average = float(match.group(1))
            reviews_count = int(re.sub(',', '', match.group(2)))
        else:
            reviews_average = None
            reviews_count = None

        business.reviews_average = reviews_average
        business.reviews_count = reviews_count
    else:
        business.reviews_average = None
        business.reviews_count = None

内容的提问来源于stack exchange,提问作者iqra naz

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.05 10:12:09