Google Map评论评分与评论数爬取异常:仅获取前4条数据
问题描述
使用Playwright爬取Google地图商家信息,计划爬取8条数据,但仅能提取前4条的评论评分(reviews_average)和评论数(reviews_count),剩余4条无法获取该类数据。
爬取代码
for index in range(min(total, len(listings))): try: listings[index].click() page.wait_for_timeout(8000) except playwright._impl._api_types.Error as e: print(f"Error clicking on listing {index + 1}: {e}") continue name_xpath = f'(//div[contains(@class, "qBF1Pd fontHeadlineSmall ")])[{index + 1}]' address_xpath = '//button[@data-item-id="address"]//div[contains(@class, "fontBodyMedium")]' website_xpath = '//a[@data-item-id="authority"]//div[contains(@class, "fontBodyMedium")]' phone_number_xpath = '//button[contains(@data-item-id, "phone:tel:")]//div[contains(@class, "fontBodyMedium")]' category_xpath = '//*[@id="QA0Szd"]/div/div/div[1]/div[3]/div/div[1]/div/div/div[2]/div[2]/div/div[1]/div[2]/div/div[2]/span/span/button' reviews_span_xpath = f'//div[{index + 1}]//span[@role="img"]' business = Business() if page.locator(name_xpath).count() > 0: business.name = page.locator(name_xpath).inner_text() else: business.name = "N/A" if page.locator(address_xpath).count() > 0: business.address = page.locator(address_xpath).inner_text() else: business.address = "N/A" if page.locator(website_xpath).count() > 0: business.website = page.locator(website_xpath).inner_text() else: business.website = "N/A" if page.locator(phone_number_xpath).count() > 0: business.phone_number = page.locator(phone_number_xpath).inner_text() else: business.phone_number = "N/A" if page.locator(category_xpath).count() > 0: business.category = page.locator(category_xpath).inner_text() else: business.category = "N/A" page.mouse.wheel(0, 20000) page.wait_for_timeout(1000) reviews_elements = listings[index].locator(reviews_span_xpath) # Move this line inside the loop if reviews_elements.count() > 0: reviews_label = reviews_elements.nth(0).get_attribute("aria-label") print(f"Reviews Label: {reviews_label}") # Process the reviews label to extract count and average # Example processing using regular expressions import re match = re.match(r'([\d.]+) stars ([\d,]+) Reviews', reviews_label) if match: reviews_average = float(match.group(1)) reviews_count = int(re.sub(',', '', match.group(2))) else: reviews_average = None reviews_count = None # Update the Business object with reviews information business.reviews_average = reviews_average business.reviews_count = reviews_count else: business.reviews_average = None business.reviews_count = None
输出结果
Total Scraped: 8 Reviews Label: 4.6 stars 117 Reviews Reviews Label: 5.0 stars 1 Reviews Reviews Label: 4.6 stars 1,308 Reviews Reviews Label: 4.9 stars 1,021 Reviews
问题分析与解决办法
核心问题
- 全局XPATH定位失效:
reviews_span_xpath用页面全局div索引定位,滚动页面后列表项DOM结构变化,索引不再对应目标商家。 - 滚动时机错误:点击商家后滚动整个页面,可能导致当前列表项DOM被移出视图或结构变更,无法定位评论元素。
- 未利用列表上下文:未基于
listings[index]的上下文定位评论元素,全局XPATH易受页面动态结构影响。
修复步骤
- 改用相对路径定位评论元素:基于当前列表项的上下文查找,避免全局索引干扰:
# 替换原reviews_span_xpath定义 reviews_span_xpath = './/span[@role="img"]' # 相对路径,仅在当前listings[index]范围内查找 - 优化等待机制:替换固定超时等待,改用Playwright的元素加载等待,确保页面就绪:
# 替换page.wait_for_timeout(8000) listings[index].click() page.wait_for_selector('//button[@data-item-id="address"]', timeout=10000) - 移除不必要的全局滚动:若评论在详情面板内,滚动详情面板而非整个页面,避免干扰列表DOM:
# 替换page.mouse.wheel(0, 20000) detail_panel = page.locator('//div[@id="QA0Szd"]//div[contains(@class, "m6QErb")]') detail_panel.scroll_into_view_if_needed() - 适配正则匹配格式:兼容不同地区的评论文本表述,避免格式不匹配导致提取失败:
# 替换原正则匹配 match = re.match(r'([\d.]+) stars? ([\d,]+) (?:Reviews?|评价)', reviews_label)
修复后的核心代码片段
for index in range(min(total, len(listings))): try: listings[index].click() # 等待详情面板加载完成 page.wait_for_selector('//button[@data-item-id="address"]', timeout=10000) except playwright._impl._api_types.Error as e: print(f"点击第{index + 1}个商家出错: {e}") continue # 其他元素定位逻辑保持不变... # 调整评论元素定位为相对路径 reviews_span_xpath = './/span[@role="img"]' reviews_elements = listings[index].locator(reviews_span_xpath) if reviews_elements.count() > 0: reviews_label = reviews_elements.nth(0).get_attribute("aria-label") print(f"评论标签: {reviews_label}") import re # 适配多语言评论文本格式 match = re.match(r'([\d.]+) stars? ([\d,]+) (?:Reviews?|评价)', reviews_label) if match: reviews_average = float(match.group(1)) reviews_count = int(re.sub(',', '', match.group(2))) else: reviews_average = None reviews_count = None business.reviews_average = reviews_average business.reviews_count = reviews_count else: business.reviews_average = None business.reviews_count = None
内容的提问来源于stack exchange,提问作者iqra naz
相关产品推荐
相关产品推荐

