You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Selenium爬取特斯拉消费者评价时XPath无效选择器报错如何解决

问题原因

你遇到的Invalid selector error报错由两个低级错误共同导致:

  1. 提取评论内容的XPath存在语法错误://*[@id="' + x +'" ]]/div[3]/p[2]/text()中id匹配结束后多写了一个多余的]闭合符
  2. Selenium的元素查找方法仅支持定位DOM元素节点,不能直接定位text()文本节点,你只需要定位到对应的p标签后读取.text属性即可拿到文本内容

另外你原代码中用户ID的获取逻辑也存在问题:当前你调用get_attribute('itemprop')拿到的是固定值author,不是实际的用户名,需要改为读取元素的.text属性。

修改后的优化代码

推荐直接遍历评论块元素使用相对XPath查找子元素,避免全局查找的匹配误差和性能问题:

import pandas as pd
from selenium import webdriver

# consumeraffairs.com站点评论爬虫类
class CarForumCrawler(): 
    def __init__(self, start_link):
        self.link_to_explore = start_link 
        self.comments = pd.DataFrame(columns = ['rating','user_id','comments'])
        # 高版本Selenium不需要手动指定chromedriver路径,会自动匹配浏览器版本
        self.driver = webdriver.Chrome()            
        self.driver.get(self.link_to_explore)
        self.driver.implicitly_wait(10)
        self.extract_data()
        self.save_data_to_file()
   
    def extract_data(self):
        # 直接获取所有评论块元素
        review_blocks = self.driver.find_elements_by_xpath("//*[contains(@id,'review-')]")
        for review in review_blocks:
            # 相对路径查找当前评论块内的评分元素
            user_rating = review.find_element_by_xpath('./div[1]/div/img')
            rating = user_rating.get_attribute('data-rating')

            # 提取用户名
            userid_element = review.find_element_by_xpath('./div[2]/div[2]/strong')
            userid = userid_element.text

            # 提取评论内容
            user_message = review.find_element_by_xpath('./div[3]/p[2]')
            comment = user_message.text

            # 存入数据框
            self.comments.loc[len(self.comments)] = [rating,userid,comment]

    def save_data_to_file(self):
        # 保存为CSV文件,指定utf-8-sig编码避免中文乱码
        self.comments.to_csv ('Tesla_rating-6.csv', index = None, header=True, encoding='utf-8-sig')
    def close_spider(self):
        self.driver.quit()

try:
    url = 'https://www.consumeraffairs.com/automotive/tesla_motors.html'
    mycrawler = CarForumCrawler(url)
    mycrawler.close_spider()
except Exception as e:
    print(e)
    raise

补充说明

如果需要爬取全量评论,需要额外添加页面滚动逻辑,触发懒加载加载更多评论内容,否则只能获取到首屏加载的评论数据。

内容的提问来源于stack exchange,提问作者Mumid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.23 22:15:03