Scrapy爬虫抓取midi-karaoke.info行为异常求助
问题求助
我在Windows 10系统上使用Python 3.11和Scrapy 2.7.1,参考Scrapy文件下载示例(原示例用于抓取nirsoft.net文件)修改脚本,目标抓取网站https://www.midi-karaoke.info,目前遇到以下问题:
- 想要抓取10万+HTML页面,但无法获取
.mid文件。该网站存在特殊情况:采用扁平化设计,拥有10万+带编号的页面名称;页面上的MIDI文件下载链接点击无反应,但在浏览器开发者工具中点击源码里的.mid链接可下载,或将页面链接后缀改为.mid(如https://www.midi-karaoke.info/21110cbd.html改为https://www.midi-karaoke.info/21110cbd.mid)也可下载。 - 修改后的脚本有时有效有时无效,相同脚本在不同时间运行结果不同。
环境信息
- 操作系统:Windows 10
- Python版本:3.11
- Scrapy版本:2.7.1
现有代码
爬虫脚本
import scrapy from scrapy.linkextractors import LinkExtractor from scrapy.spiders import CrawlSpider, Rule from webcrawler.items import WebcrawlerItem # import C:\..\scrapy\webcrawler\webcrawler\items.py class WebcrawlSpider(CrawlSpider): name = 'webcrawl' allowed_domains = ['www.midi-karaoke.info'] start_urls = ['https://www.midi-karaoke.info'] # Redirections vermeiden ? custom_settings = {'REDIRECT_ENABLED': False} handle_httpstatus_list = [302, 301] rules = ( Rule(LinkExtractor(allow=r'/'), callback='parse_item', follow=True), # webseite befindet sich nur in '/' Rule(LinkExtractor(allow=(r'/'), deny_extensions=[], restrict_xpaths=('//a[@href]')), callback="parse_items", follow= True), # extrahiere 'href' links Rule(LinkExtractor(allow=(r'/'), restrict_xpaths=('//a[@class="MIDI"]',)), callback="parse_items", follow= True), # href links die uns interessieren befinden sich in class='MIDI' ) def parse_item(self, response): file_url = response.css('.downloadline::attr(href)').getall() # hole alle gefundenen Seiten file_url = response.urljoin(file_url) file_extension = file_url.split('.')[-1] # filtere links nach Dateien mir Extension (Optional) if file_extension not in ('mid' , 'html', 'zip'): return #if '.ru.' in file_url or '.en.' in file_url: # return item = WebcrawlerItem() item['file_urls'] = [file_url] item['original_file_name'] = file_url.split('/')[-1] yield item
settings.py
# Scrapy settings for webcrawler project BOT_NAME = 'webcrawler' SPIDER_MODULES = ['webcrawler.spiders'] NEWSPIDER_MODULE = 'webcrawler.spiders' DUPEFILTER_DEBUG = False REDIRECT_ENABLED = False ROBOTSTXT_OBEY = True ITEM_PIPELINES = { 'webcrawler.pipelines.WebcrawlerPipeline': 1, } FILES_STORE = r"C:\Users\wiwa53\scrapy\webcrawler\downloads" REQUEST_FINGERPRINTER_IMPLEMENTATION = '2.7' TWISTED_REACTOR = 'twisted.internet.asyncioreactor.AsyncioSelectorReactor' USER_AGENT = 'Mozilla/5.0 (iPad; CPU OS 12_2 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Mobile/15E148'
items.py
# Define here the models for your scraped items import scrapy class WebcrawlerItem(scrapy.Item): file_urls = scrapy.Field() original_file_name = scrapy.Field() files = scrapy.Field
pipelines.py
# Define your item pipelines here from scrapy.pipelines.files import FilesPipeline class WebcrawlerPipeline(FilesPipeline): def file_path(self, request, response=None, info=None, *, item=None): file_name: str = request.url.split("/")[-1] #print(file_name) return file_name
内容的提问来源于stack exchange,提问作者wiwa53
相关产品推荐
相关产品推荐

