使用Scrapy框架爬取电商网站产品信息遇脚本问题求助
Scrapy脚本问题修复方案
原有代码核心错误
- 提取商品链接时使用
get()方法仅能获取第一条匹配的链接,后续循环实际是遍历单个字符串的每个字符,逻辑完全错误,需替换为getall()获取全部商品链接。 - 拼接得到完整商品链接后,错误调用字符串对象的
css()方法,字符串无该属性会直接触发运行报错,直接返回拼接完成的product_link即可。 - 评论区选择器
.Natsob与目标站点实际结构不匹配,无法提取到评论内容;同时标题、价格提取结果存在多余空白字符,需要做清洗处理。
修正后的完整代码
import scrapy class DressSpider(scrapy.Spider): name = 'dress' allowed_domains = ['savedbythedress.com'] start_urls = ['https://savedbythedress.com/collections/maternity-tops'] def parse(self, response): domain = "https://savedbythedress.com" # 改为getall()获取所有商品链接 link_products = response.css('div.product-info-inner a::attr(href)').getall() for link in link_products: product_link = domain + link # 直接返回拼接好的链接,不要调用css方法 yield{ 'product_link': product_link, } yield scrapy.Request(url=product_link, callback=self.parse_contents) def parse_contents(self, response): yield{ 'product_title' : response.css('.sbtd-product-title ::text').get().strip(), 'product_price' : response.css('.product-price ::text').get().strip(), # 修正评论选择器,适配目标站点的评论结构 'product_review' : [rev.strip() for rev in response.css('.jdgm-rev__body ::text').getall() if rev.strip()] }
运行注意事项
- 若运行时出现403拦截,可在Scrapy项目的
settings.py中添加合法的请求头标识,示例配置:USER_AGENT = 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' - 若站点开启了动态渲染,可搭配scrapy-splash或者playwright组件加载动态内容,保证评论数据完整提取。
内容的提问来源于stack exchange,提问作者Thinesh
相关产品推荐
相关产品推荐

