Scrapy未发起请求求助:测试爬取Youtube无Get请求
Hey there! Let's break down why your Scrapy spider isn't making any requests and fix it step by step.
核心问题:方法名拼写错误
The biggest issue here is a typo in the method name. Scrapy expects spiders to define a start_requests() method (note the plural "s" at the end) to initiate requests. You've written start_request() (singular), so Scrapy doesn't recognize this method and never triggers any requests—that's exactly why you see the spider open and immediately close without any GET request logs.
其他需要修正的细节
Even after fixing the method name, there are a few other adjustments to make your spider run smoothly:
处理标题的列表格式:
response.css('title::text').extract()returns a list. Trying to write a list directly to a file will throw an error. Useextract_first()(orget()in newer Scrapy versions) to get a single string value:title = response.css('title::text').extract_first() # Or for Scrapy 1.5+: title = response.css('title::text').get()ROBOTSTXT_OBEY 设置: Your logs show
ROBOTSTXT_OBEY: True. YouTube's robots.txt file likely blocks crawlers, so even with the fixed method name, your request might be rejected. For testing purposes, you can temporarily setROBOTSTXT_OBEY = Falsein yoursettings.pyfile.文件编码: When writing to a file, specify the encoding to avoid potential garbled text:
with open('informacao', 'w', encoding='utf-8') as f: f.write(title)
修正后的完整代码
Here's the polished version of your spider:
import scrapy class TesteSpider(scrapy.Spider): name = "teste" def start_requests(self): # Fixed method name url = 'http://www.youtube.com' yield scrapy.Request(url, self.parse) def parse(self, response): title = response.css('title::text').extract_first() # Get single string with open('informacao', 'w', encoding='utf-8') as f: f.write(title) self.log('saved file successfully')
为什么 Scrapy Shell 能正常工作?
The scrapy shell command directly sends the request you specify—it doesn't rely on the start_requests() method of a spider. That's why it works even when your spider has the method name typo.
内容的提问来源于stack exchange,提问作者Silvio Machado

