Play Store评论爬虫仅获首屏数据求助(附Python代码)
解决Google Play Store评论爬取仅获取首屏数据的问题
Google Play的评论在点击“显示全部评论”后,后续内容是通过JavaScript动态加载的。requests库只能获取页面初始的静态HTML,无法捕获异步加载的评论数据,这就是你的脚本只能拿到首屏评论的原因。
以下提供两种可行的解决方案:
方案1:使用Selenium模拟浏览器动态加载(适合快速实现)
Selenium可以模拟真实浏览器操作,自动触发评论加载、滚动页面获取所有数据,再解析完整的页面内容。
步骤与代码实现
- 先安装依赖:
pip install selenium beautifulsoup4
- 下载对应浏览器的驱动(如ChromeDriver,需与浏览器版本匹配),并配置系统环境变量或指定驱动路径。
修改后的爬取代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup import time def scrape_playstore_all_reviews(url): reviews_list = [] # 初始化Chrome浏览器(若驱动未配置环境变量,需指定executable_path参数) driver = webdriver.Chrome() driver.get(url) try: # 等待并点击"显示全部评论"按钮(适配葡萄牙语页面文本) show_all_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, '//button[contains(text(), "Mostrar todos os comentários")]')) ) show_all_btn.click() # 循环滚动页面,加载所有评论 last_height = driver.execute_script("return document.body.scrollHeight") while True: driver.execute_script("window.scrollTo(0, document.body.scrollHeight);") time.sleep(2) # 等待评论加载 new_height = driver.execute_script("return document.body.scrollHeight") if new_height == last_height: break last_height = new_height # 解析加载完成后的页面 soup = BeautifulSoup(driver.page_source, 'html.parser') app_title = soup.find('h1', class_='Fd93Bb').find('span').text # 提取所有评论数据 reviews = soup.find_all('div', class_='EGFGHd') for review in reviews: user_name = review.find('div', class_='X5PpBb').text if review.find('div', class_='X5PpBb') else None comment_text = review.find('div', class_='h3YV2d').text if review.find('div', class_='h3YV2d') else None helpful_count = review.find('div', class_='AJTPZc').text if review.find('div', class_='AJTPZc') else None post_date = review.find('span', class_='bp9Aid').text if review.find('span', class_='bp9Aid') else None star_rating = review.find('div', class_='iXRFPc')['aria-label'] if review.find('div', class_='iXRFPc') else None review_data = { 'Título do aplicativo': app_title, 'Nome do usuário': user_name, 'Texto do comentário': comment_text, 'Qtd apoio ao comentário': helpful_count, 'Data da postagem': post_date, 'Qtd estrelas': star_rating } reviews_list.append(review_data) print(review_data) finally: driver.quit() # 确保关闭浏览器,释放资源 return reviews_list url = 'https://play.google.com/store/apps/details?id=ru.zengalt.simpler&hl=pt_BR' reviews = scrape_playstore_all_reviews(url)
代码说明
- 使用
WebDriverWait等待元素加载,避免因页面渲染延迟导致的元素定位失败 - 通过JavaScript滚动页面,循环判断页面高度是否停止变化,确保所有评论加载完成
- 提取字段时增加空值判断,避免单个评论字段缺失引发报错
方案2:调用Google Play官方评论API(进阶)
如果需要大规模、稳定爬取评论,可以使用Google Play Developer API的评论接口。但需要提前完成:
- 在Google Cloud Console创建项目,启用Google Play Developer API
- 获取OAuth 2.0认证凭据
- 确保拥有目标应用的开发者权限(或使用公开数据接口)
示例请求代码:
import requests API_KEY = '你的API密钥' APP_PACKAGE = 'ru.zengalt.simpler' url = f'https://play.googleapis.com/admin/v1/reviews/{APP_PACKAGE}?key={API_KEY}' response = requests.get(url) reviews_data = response.json()
该方案有请求配额限制,适合正式研究项目使用,个人快速验证用Selenium更便捷。
反爬注意事项
- 避免短时间内频繁请求,建议适当延长等待时间
- Google Play页面元素类名可能随版本更新变化,若后续爬取失败,需重新检查元素定位规则
内容的提问来源于stack exchange,提问作者Iasmim Godoy
相关产品推荐
相关产品推荐

