Python脚本爬取Trustpilot评论异常:仅生成CSV表头无数据
问题
本人是Python及编程新手,编写了一段爬取Trustpilot客户评论的脚本,在Google Bard中测试可返回结果,但在Mac的PyCharm CE中运行时,仅生成带正确表头的CSV文件,无评论数据。本地运行无报错,已安装Python 3.12及所有所需模块,求问原因及解决办法。
用户脚本代码:
from selenium import webdriver from selenium.webdriver.common.by import By from bs4 import BeautifulSoup import csv import datetime # Create a new Selenium webdriver instance driver = webdriver.Chrome() # Navigate to the given page driver.get("https://uk.trustpilot.com/review/www.whsmith.co.uk") # Wait for the page to load driver.implicitly_wait(10) # Get the HTML source code of the page html = driver.page_source # Create a BeautifulSoup object from the HTML source code soup = BeautifulSoup(html, "html.parser") # Extract all of the reviews from the page reviews = soup.findAll("div", class_="review") # Create a new CSV file to store the reviews with open("whsmith_reviews.csv", "w", newline="") as f: writer = csv.writer(f) # Write the header row writer.writerow(["Review Title", "Review Text", "Rating", "Review Date"]) # Iterate over the reviews and write them to the CSV file for review in reviews: title = review.find("h2", class_="review-title").text text = review.find("p", class_="review-text").text rating = review.find("span", class_="review-rating").text date_str = review.find("span", class_="review-date").text date = datetime.datetime.strptime(date_str, "%d %b %Y") # Add the review to the CSV file writer.writerow([title, text, rating, date]) # Close the Selenium webdriver instance driver.quit()
原因分析
- CSS选择器失效:Trustpilot页面元素的类名是动态生成的,本地加载的页面结构和Google Bard测试环境的页面结构存在差异,导致代码中
review等类名无法匹配到实际评论元素。 - 页面加载不充分:隐式等待10秒的时长不足以适配本地网络速度,评论内容还未渲染完成就已获取页面源码。
- 反爬机制拦截:Trustpilot可能识别出本地Selenium的自动化请求,限制了评论内容的加载,仅返回页面框架。
- ChromeDriver版本不兼容:本地Chrome浏览器版本与ChromeDriver版本不匹配,导致页面渲染异常,无法获取完整的HTML内容。
解决办法
1. 修正元素选择器
打开目标页面,按F12调出开发者工具,查看评论容器及子元素的实际类名,替换代码中的选择器。比如当前Trustpilot的评论容器类名可能为styles_reviewCard__9HxJJ,子元素类名也需同步更新:
# 替换评论容器选择器 reviews = soup.findAll("div", class_="styles_reviewCard__9HxJJ") # 替换子元素选择器示例 title = review.find("h2", class_="styles_reviewTitle__04VGJ").text.strip() text = review.find("p", class_="styles_reviewContent__0Q2Tg").text.strip() rating = review.find("div", class_="styles_starRating__4rrcf")["data-rating"] date_str = review.find("time")["datetime"].split("T")[0]
2. 优化等待机制
改用显式等待,确保评论元素加载完成后再获取页面源码:
from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC # 等待评论元素出现,超时时间设为20秒 WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.CLASS_NAME, "styles_reviewCard__9HxJJ")) ) html = driver.page_source
3. 绕过反爬检测
为ChromeDriver添加参数,伪装成普通浏览器请求:
options = webdriver.ChromeOptions() options.add_argument("--headless=new") # 无头模式,可根据需求关闭 options.add_argument("--disable-blink-features=AutomationControlled") options.add_argument("user-agent=Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36") driver = webdriver.Chrome(options=options)
若存在滚动加载,可添加滚动操作触发评论加载:
driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
4. 匹配ChromeDriver版本
查看Chrome浏览器版本(Chrome菜单→关于Google Chrome),下载对应版本的ChromeDriver。Mac用户可通过Homebrew快速安装匹配版本:
brew install chromedriver
内容的提问来源于stack exchange,提问作者Justin Neale
相关产品推荐
相关产品推荐

