Scrapy爬虫无法爬取数据,VSCode报'yield' outside function错误
Scrapy爬虫问题修复与代码优化
1. 语法错误(SyntaxError: 'yield' outside function)修复
你的parse方法缩进错误,未作为multiSpider类的成员方法,导致被定义为全局函数,代码结构不符合Scrapy要求。将parse方法缩进至类内部,与name、start_urls同级即可解决该问题。
2. Scrapy核心属性修正
Scrapy爬虫类中起始URL列表的属性名必须是**start_urls**(复数形式),而非start_url,否则Scrapy无法识别并加载起始页面。
3. CSS选择器修正
原选择器存在语法错误和匹配逻辑问题,修正如下:
- 回答数选择器:将
threadstats td alt :: text改为.threadstats td .threadreplies::text(目标页面中回答数位于.threadstats下的.threadreplies元素内) - 作者选择器:将
a.username offline popupctrl :: text改为a.username.offline.popupctrl::text(多class选择器需用.连接,空格代表后代选择器,不符合当前匹配需求) - 分页链接选择器:将
span.selected pageitem a::attr(href)改为span.selected + .pageitem a::attr(href)(匹配当前选中页后的下一页按钮)
4. 分页逻辑修正
原代码中response.urljoin('next_page')错误传入字符串'next_page',应传入获取到的分页链接变量next_page,否则会生成无效URL。
完整修正后的代码
import scrapy class multiSpider(scrapy.Spider): name='multiple' start_urls = [ 'https://forum.moshaver.co/f232/', 'https://forum.moshaver.co/f233/', 'https://forum.moshaver.co/f241/', 'https://forum.moshaver.co/f231/', ] def parse(self, response): for data in response.css('h3.threadtitle'): yield { 'title': data.css('a::text').get().strip(), 'answers': data.css('.threadstats td .threadreplies::text').get(), 'writer': data.css('a.username.offline.popupctrl::text').get(), 'date_time': data.css('span.label a::text').get(), } next_page = response.css('span.selected + .pageitem a::attr(href)').get() if next_page: next_page = response.urljoin(next_page) yield scrapy.Request(url=next_page, callback=self.parse)
内容的提问来源于stack exchange,提问作者zahra jamalou
相关产品推荐
相关产品推荐

