You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup+Selenium无法检测到review_body等Div元素的问题

问题

我在解析某专辑评论页面时遇到问题:本地下载HTML文件后,能正常获取class为review_body的div元素,但用Selenium访问页面时却检测不到这些元素。我知道可能需要等待JavaScript加载,但尝试相关操作后仍未解决。

我的代码片段
for album_url in albums:
    print(f"Processing album: {album_url}")
    has_reviews = True  # 标记是否有评论
    for page in range(1, 100):  # 最多爬100页
        try:
            url = f"{album_url}{page}/"
            print(f"正在爬取页面: {url}")
            driver.get(url)

            # 用BeautifulSoup解析页面源码
            soup = BeautifulSoup(driver.page_source, 'html.parser')

            # 提取评论
            header = soup.find_all('div', class_='review_header')
            print(header)
            reviews = soup.find_all('div', class_='review_body')
            print(reviews)

            if not reviews:
                print(f"第{page}页未找到评论,切换到下一张专辑。")
                has_reviews = False
                break
可行解决方案
  • 用显式等待替代立即解析:driver.get()仅等待页面初始加载,评论可能是异步渲染的,需等待目标元素出现后再解析:
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.common.exceptions import TimeoutException

# ... 循环内的代码 ...
driver.get(url)

try:
    # 最长等待10秒,直到至少一个review_body元素加载完成
    WebDriverWait(driver, 10).until(
        EC.presence_of_element_located((By.CLASS_NAME, "review_body"))
    )
except TimeoutException:
    print("超时未加载到评论元素,跳过当前页")
    continue

# 此时再解析页面
soup = BeautifulSoup(driver.page_source, 'html.parser')
reviews = soup.find_all('div', class_='review_body')
  • 检查是否需要滚动触发加载:部分页面需要滚动到底部才会加载评论内容,可添加滚动操作:
driver.get(url)
# 执行JS滚动到页面底部
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
# 等待加载(或用显式等待替代sleep)
import time
time.sleep(2)
  • 规避反爬检测:该网站可能识别Selenium自动化特征,可修改浏览器参数隐藏自动化标识:
from selenium.webdriver.chrome.options import Options

options = Options()
# 禁用自动化提示
options.add_argument("--disable-blink-features=AutomationControlled")
# 设置自定义UA
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 可选:无头模式
options.add_argument("--headless=new")

driver = webdriver.Chrome(options=options)
  • 直接用Selenium定位元素验证:跳过BeautifulSoup,直接用Selenium查找元素,确认是否真的未加载:
driver.get(url)
# 等待后查找元素
reviews = driver.find_elements(By.CLASS_NAME, "review_body")
print(f"找到{len(reviews)}条评论")

内容的提问来源于stack exchange,提问作者Nate

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.16 05:13:10