You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Beautiful Soup find_all无法提取谷歌Play Store页面内容求助

解决Google Play Store评论爬取失败的问题

Hey Neil, I’ve run into this exact issue before—let’s break down why your current approach isn’t working and fix it.

为什么爬Newegg能用,但Google Play不行?

Newegg提供的是静态HTML——所有产品数据都包含在uReq()获取的初始页面响应里。但Google Play Store用的是动态JavaScript渲染:评论(以及大部分应用详情)并不在初始HTML中。当你用urllib抓取页面时,只能拿到页面的空框架,那些通过JS后续加载的评论内容根本没被获取到。这就是为什么page_soup.find_all("div",{"class":"zc7KVe"})返回空列表——这些元素在你解析的原始HTML里根本不存在。

两种可行的解决方案

方案1:用Selenium模拟浏览器加载(适合需要完整页面渲染的场景)

Selenium会启动一个真实浏览器(或无头浏览器)执行JavaScript,这样你就能拿到包含评论的完整渲染页面。调整代码如下:

首先安装Selenium并下载对应Chrome版本的ChromeDriver:

pip install selenium

然后使用以下代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from bs4 import BeautifulSoup as soup

# 初始化无头Chrome(移除--headless可以看到浏览器窗口)
options = webdriver.ChromeOptions()
options.add_argument("--headless=new")
driver = webdriver.Chrome(options=options)

my_url = "https://play.google.com/store/apps/details?id=com.disney.disneyplus&hl=de&showAllReviews=true"
driver.get(my_url)

# 等待评论加载(可根据网络情况调整超时时间)
try:
    # 等待至少一个评论容器出现
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "zc7KVe"))
    )
finally:
    # 获取完整渲染后的HTML
    page_html = driver.page_source
    driver.quit()

# 用BeautifulSoup解析
page_soup = soup(page_html, "html.parser")
reviews = page_soup.find_all("div", {"class": "zc7KVe"})

# 打印第二条评论(和你原代码逻辑一致)
print(reviews[1])

方案2:用专门的Google Play爬取库(更简单可靠)

不用自己写爬虫,直接用google-play-scraper库——它直接调用Google的内部API,无需处理动态渲染或HTML解析的问题。

安装库:

pip install google-play-scraper

然后用几行代码就能获取评论:

from google_play_scraper import reviews

# 获取Disney+的德语区评论
result, continuation_token = reviews(
    "com.disney.disneyplus",
    lang="de",
    country="de",
    count=100,  # 要获取的评论数量
    sort=reviews.SORT_MOST_RELEVANT,
)

# result里的每个元素都是包含评论详情的字典
for review in result[:2]:  # 打印前2条评论
    print("评分:", review["score"])
    print("评论内容:", review["content"])
    print("---")

重要注意事项

  • 请求频率限制:如果请求太频繁,Google会封禁你的IP。用Selenium时可以加time.sleep()延迟,用google-play-scraper时可以用continuation_token逐步获取更多评论。
  • 区域设置:确保lang和country和你目标页面一致(你原URL用了hl=de,所以保持lang="de", country="de")。
  • 无头模式:Selenium的--headless参数可以让浏览器在后台运行,更适合爬虫脚本。

内容的提问来源于stack exchange,提问作者O'Neil Murray

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 07:28:01