Beautiful Soup find_all无法提取谷歌Play Store页面内容求助
解决Google Play Store评论爬取失败的问题
Hey Neil, I’ve run into this exact issue before—let’s break down why your current approach isn’t working and fix it.
为什么爬Newegg能用,但Google Play不行?
Newegg提供的是静态HTML——所有产品数据都包含在uReq()获取的初始页面响应里。但Google Play Store用的是动态JavaScript渲染:评论(以及大部分应用详情)并不在初始HTML中。当你用urllib抓取页面时,只能拿到页面的空框架,那些通过JS后续加载的评论内容根本没被获取到。这就是为什么page_soup.find_all("div",{"class":"zc7KVe"})返回空列表——这些元素在你解析的原始HTML里根本不存在。
两种可行的解决方案
方案1:用Selenium模拟浏览器加载(适合需要完整页面渲染的场景)
Selenium会启动一个真实浏览器(或无头浏览器)执行JavaScript,这样你就能拿到包含评论的完整渲染页面。调整代码如下:
首先安装Selenium并下载对应Chrome版本的ChromeDriver:
pip install selenium
然后使用以下代码:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup as soup # 初始化无头Chrome(移除--headless可以看到浏览器窗口) options = webdriver.ChromeOptions() options.add_argument("--headless=new") driver = webdriver.Chrome(options=options) my_url = "https://play.google.com/store/apps/details?id=com.disney.disneyplus&hl=de&showAllReviews=true" driver.get(my_url) # 等待评论加载(可根据网络情况调整超时时间) try: # 等待至少一个评论容器出现 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CLASS_NAME, "zc7KVe")) ) finally: # 获取完整渲染后的HTML page_html = driver.page_source driver.quit() # 用BeautifulSoup解析 page_soup = soup(page_html, "html.parser") reviews = page_soup.find_all("div", {"class": "zc7KVe"}) # 打印第二条评论(和你原代码逻辑一致) print(reviews[1])
方案2:用专门的Google Play爬取库(更简单可靠)
不用自己写爬虫,直接用google-play-scraper库——它直接调用Google的内部API,无需处理动态渲染或HTML解析的问题。
安装库:
pip install google-play-scraper
然后用几行代码就能获取评论:
from google_play_scraper import reviews # 获取Disney+的德语区评论 result, continuation_token = reviews( "com.disney.disneyplus", lang="de", country="de", count=100, # 要获取的评论数量 sort=reviews.SORT_MOST_RELEVANT, ) # result里的每个元素都是包含评论详情的字典 for review in result[:2]: # 打印前2条评论 print("评分:", review["score"]) print("评论内容:", review["content"]) print("---")
重要注意事项
- 请求频率限制:如果请求太频繁,Google会封禁你的IP。用Selenium时可以加
time.sleep()延迟,用google-play-scraper时可以用continuation_token逐步获取更多评论。 - 区域设置:确保
lang和country和你目标页面一致(你原URL用了hl=de,所以保持lang="de", country="de")。 - 无头模式:Selenium的
--headless参数可以让浏览器在后台运行,更适合爬虫脚本。
内容的提问来源于stack exchange,提问作者O'Neil Murray
相关产品推荐
相关产品推荐

