You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python脚本爬取Trustpilot评论异常:仅生成CSV表头无数据

问题

本人是Python及编程新手,编写了一段爬取Trustpilot客户评论的脚本,在Google Bard中测试可返回结果,但在Mac的PyCharm CE中运行时,仅生成带正确表头的CSV文件,无评论数据。本地运行无报错,已安装Python 3.12及所有所需模块,求问原因及解决办法。

用户脚本代码:

from selenium import webdriver
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup
import csv
import datetime

# Create a new Selenium webdriver instance
driver = webdriver.Chrome()

# Navigate to the given page
driver.get("https://uk.trustpilot.com/review/www.whsmith.co.uk")

# Wait for the page to load
driver.implicitly_wait(10)

# Get the HTML source code of the page
html = driver.page_source

# Create a BeautifulSoup object from the HTML source code
soup = BeautifulSoup(html, "html.parser")

# Extract all of the reviews from the page
reviews = soup.findAll("div", class_="review")

# Create a new CSV file to store the reviews
with open("whsmith_reviews.csv", "w", newline="") as f:
    writer = csv.writer(f)

    # Write the header row
    writer.writerow(["Review Title", "Review Text", "Rating", "Review Date"])

    # Iterate over the reviews and write them to the CSV file
    for review in reviews:
        title = review.find("h2", class_="review-title").text
        text = review.find("p", class_="review-text").text
        rating = review.find("span", class_="review-rating").text
        date_str = review.find("span", class_="review-date").text
        date = datetime.datetime.strptime(date_str, "%d %b %Y")

        # Add the review to the CSV file
        writer.writerow([title, text, rating, date])

# Close the Selenium webdriver instance
driver.quit()

原因分析

  • CSS选择器失效:Trustpilot页面元素的类名是动态生成的,本地加载的页面结构和Google Bard测试环境的页面结构存在差异,导致代码中review等类名无法匹配到实际评论元素。
  • 页面加载不充分:隐式等待10秒的时长不足以适配本地网络速度,评论内容还未渲染完成就已获取页面源码。
  • 反爬机制拦截:Trustpilot可能识别出本地Selenium的自动化请求,限制了评论内容的加载,仅返回页面框架。
  • ChromeDriver版本不兼容:本地Chrome浏览器版本与ChromeDriver版本不匹配,导致页面渲染异常,无法获取完整的HTML内容。

解决办法

1. 修正元素选择器

打开目标页面,按F12调出开发者工具,查看评论容器及子元素的实际类名,替换代码中的选择器。比如当前Trustpilot的评论容器类名可能为styles_reviewCard__9HxJJ,子元素类名也需同步更新:

# 替换评论容器选择器
reviews = soup.findAll("div", class_="styles_reviewCard__9HxJJ")

# 替换子元素选择器示例
title = review.find("h2", class_="styles_reviewTitle__04VGJ").text.strip()
text = review.find("p", class_="styles_reviewContent__0Q2Tg").text.strip()
rating = review.find("div", class_="styles_starRating__4rrcf")["data-rating"]
date_str = review.find("time")["datetime"].split("T")[0]

2. 优化等待机制

改用显式等待,确保评论元素加载完成后再获取页面源码:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 等待评论元素出现,超时时间设为20秒
WebDriverWait(driver, 20).until(
    EC.presence_of_element_located((By.CLASS_NAME, "styles_reviewCard__9HxJJ"))
)

html = driver.page_source

3. 绕过反爬检测

为ChromeDriver添加参数,伪装成普通浏览器请求:

options = webdriver.ChromeOptions()
options.add_argument("--headless=new")  # 无头模式,可根据需求关闭
options.add_argument("--disable-blink-features=AutomationControlled")
options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36")

driver = webdriver.Chrome(options=options)

若存在滚动加载,可添加滚动操作触发评论加载:

driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")

4. 匹配ChromeDriver版本

查看Chrome浏览器版本(Chrome菜单→关于Google Chrome),下载对应版本的ChromeDriver。Mac用户可通过Homebrew快速安装匹配版本:

brew install chromedriver

内容的提问来源于stack exchange,提问作者Justin Neale

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 19:12:23