Scrapy CSS选择器无法抓取目标网站卡片信息问题求助
问题解决:Scrapy无法定位BanChile网站卡片内容
核心问题分析
- CSS选择器语法错误:你写的
.new-beneficios-card-title d-flex::text写法错误,d-flex是元素类名,多个类名需用.连接,正确格式应为.new-beneficios-card-title.d-flex::text;实际页面中,标题元素仅用.new-beneficios-card-title即可定位,多余类名反而可能干扰匹配。 - 动态内容加载问题:该网站卡片内容由JavaScript动态渲染,Scrapy默认请求只能获取初始HTML,无法拿到JS加载后的内容,
time.sleep()在此完全无效——Scrapy是异步架构,sleep只会阻塞进程,不会等待页面渲染。 - Item提取逻辑不合理:你直接将所有title和summary存入单个Item,正确做法是遍历每个卡片,为每个卡片生成独立Item。
解决方案
1. 修正CSS选择器
目标元素的正确选择器:
- 卡片标题:
.new-beneficios-card-title::text - 卡片副标题:
.new-beneficios-card-subtitle::text
2. 处理动态加载内容
使用Scrapy结合Playwright处理JS渲染页面:
- 安装依赖:
pip install scrapy-playwright - 在
settings.py中配置Playwright下载器中间件:
DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } DOWNLOADER_MIDDLEWARES = { "scrapy.downloadermiddlewares.useragent.UserAgentMiddleware": None, "scrapy_playwright.middleware.PlaywrightMiddleware": 543, } PLAYWRIGHT_LAUNCH_OPTIONS = { "headless": True, # 改为False可查看浏览器窗口 "timeout": 10000, }
3. 修正爬虫代码
import scrapy from ..items import PracticescraperItem class BanChileSpider(scrapy.Spider): name = 'banchile' start_urls = [ 'https://portales.bancochile.cl/personas/beneficios?categoria=marcas' ] def start_requests(self): for url in self.start_urls: yield scrapy.Request( url, meta={"playwright": True, "playwright_include_page": True}, callback=self.parse ) async def parse(self, response): # 等待卡片元素加载完成 await response.meta["playwright_page"].wait_for_selector('.new-beneficios-card', timeout=10000) # 遍历所有卡片 for card in response.css('.new-beneficios-card'): items = PracticescraperItem() # 提取并清洗内容,避免空值 title = card.css('.new-beneficios-card-title::text').get() summary = card.css('.new-beneficios-card-subtitle::text').get() items['title'] = title.strip() if title else None items['summary'] = summary.strip() if summary else None yield items # 关闭Playwright页面 await response.meta["playwright_page"].close()
额外说明
- 避免在Scrapy中使用
time.sleep(),如需等待特定元素,优先用Playwright的wait_for_selector()方法。 - 确保
PracticescraperItem已正确定义title和summary字段。
内容的提问来源于stack exchange,提问作者Felipe Levinir
相关产品推荐
相关产品推荐

