You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫抓取midi-karaoke.info行为异常求助

问题求助

我在Windows 10系统上使用Python 3.11和Scrapy 2.7.1,参考Scrapy文件下载示例(原示例用于抓取nirsoft.net文件)修改脚本,目标抓取网站https://www.midi-karaoke.info,目前遇到以下问题:

  1. 想要抓取10万+HTML页面,但无法获取.mid文件。该网站存在特殊情况:采用扁平化设计,拥有10万+带编号的页面名称;页面上的MIDI文件下载链接点击无反应,但在浏览器开发者工具中点击源码里的.mid链接可下载,或将页面链接后缀改为.mid(如https://www.midi-karaoke.info/21110cbd.html改为https://www.midi-karaoke.info/21110cbd.mid)也可下载。
  2. 修改后的脚本有时有效有时无效,相同脚本在不同时间运行结果不同。
环境信息
  • 操作系统:Windows 10
  • Python版本:3.11
  • Scrapy版本:2.7.1
现有代码

爬虫脚本

import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule
from webcrawler.items import WebcrawlerItem # import C:\..\scrapy\webcrawler\webcrawler\items.py

class WebcrawlSpider(CrawlSpider):
    name = 'webcrawl'
    allowed_domains = ['www.midi-karaoke.info']
    start_urls = ['https://www.midi-karaoke.info']

    # Redirections vermeiden ?
    custom_settings = {'REDIRECT_ENABLED': False}
    handle_httpstatus_list = [302, 301]

    rules = (
        Rule(LinkExtractor(allow=r'/'), callback='parse_item', follow=True),
        # webseite befindet sich nur in '/'
        Rule(LinkExtractor(allow=(r'/'), deny_extensions=[], restrict_xpaths=('//a[@href]')), callback="parse_items", follow= True),
        # extrahiere 'href' links
        Rule(LinkExtractor(allow=(r'/'), restrict_xpaths=('//a[@class="MIDI"]',)), callback="parse_items", follow= True),
        # href links die uns interessieren befinden sich in class='MIDI'   
    )

    def parse_item(self, response):
        file_url = response.css('.downloadline::attr(href)').getall() # hole alle gefundenen Seiten
        file_url = response.urljoin(file_url)       
        file_extension = file_url.split('.')[-1]
        # filtere links nach Dateien mir Extension (Optional)
        if file_extension not in ('mid' , 'html', 'zip'): 
            return
        #if '.ru.' in file_url or '.en.' in file_url:
         #   return
        item = WebcrawlerItem()
        item['file_urls'] = [file_url]
        item['original_file_name'] = file_url.split('/')[-1]
        yield item

settings.py

# Scrapy settings for webcrawler project

BOT_NAME = 'webcrawler'
SPIDER_MODULES = ['webcrawler.spiders']
NEWSPIDER_MODULE = 'webcrawler.spiders'
DUPEFILTER_DEBUG = False
REDIRECT_ENABLED = False 
ROBOTSTXT_OBEY = True
ITEM_PIPELINES = {
    'webcrawler.pipelines.WebcrawlerPipeline': 1,
}
FILES_STORE = r"C:\Users\wiwa53\scrapy\webcrawler\downloads"
REQUEST_FINGERPRINTER_IMPLEMENTATION = '2.7'
TWISTED_REACTOR = 'twisted.internet.asyncioreactor.AsyncioSelectorReactor'

USER_AGENT = 'Mozilla/5.0 (iPad; CPU OS 12_2 like Mac OS X) AppleWebKit/605.1.15 (KHTML, like Gecko) Mobile/15E148'

items.py

# Define here the models for your scraped items
import scrapy

class WebcrawlerItem(scrapy.Item):
    file_urls = scrapy.Field()
    original_file_name = scrapy.Field()
    files = scrapy.Field

pipelines.py

# Define your item pipelines here
from scrapy.pipelines.files import FilesPipeline

class WebcrawlerPipeline(FilesPipeline):
    def file_path(self, request, response=None, info=None, *, item=None):        
        file_name: str = request.url.split("/")[-1]
        #print(file_name)
        return file_name

内容的提问来源于stack exchange,提问作者wiwa53

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.10 12:40:12