You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Scrapy的start_requests中用循环返回多个scrapy.Request避免重复代码?

Fixing the Loop in Scrapy's start_requests Method

Hey there! I get it—repeating scrapy.Request lines is tedious, and it makes total sense to use a loop to generate all your requests cleanly. Let's break down why your initial loop attempt might have failed, and get you set up with scalable, DRY (Don't Repeat Yourself) code.

The Core Issue

Your current code works, but manually listing three scrapy.Request objects isn't ideal. If you tried a loop before and only got the first result, chances are you either returned a single request instead of an iterable, or your loop didn't properly pair URLs with their matching callback functions.

Scrapy’s start_requests method can act as a generator—using yield for each request aligns perfectly with how Scrapy handles asynchronous tasks, and it’s super readable. Here’s how to rewrite your code:

class SiteFetching(scrapy.Spider):
    name = 'Site'
    
    def start_requests(self):
        # Pair each URL directly with its corresponding callback
        url_callback_pairs = [
            ('https://www.rev.com/freelancers/transcription', self.parse_transcription),
            ('https://www.rev.com/freelancers/captions', self.parse_caption),
            ('https://www.rev.com/freelancers/subtitles', self.parse_subtitles)
        ]
        
        # Loop through each pair and yield the request
        for url, callback in url_callback_pairs:
            yield scrapy.Request(url, callback=callback)

Solution 2: Keep Your Original Dictionary (If You Prefer)

If you want to retain your links dictionary, just use zip() to pair its values with your callback list correctly:

class SiteFetching(scrapy.Spider):
    name = 'Site'
    
    def start_requests(self):
        links = {
            'transcription_page': 'https://www.rev.com/freelancers/transcription', 
            'captions_page': 'https://www.rev.com/freelancers/captions', 
            'subtitles_page': 'https://www.rev.com/freelancers/subtitles'
        }
        callbacks = [self.parse_transcription, self.parse_caption, self.parse_subtitles]
        
        # Zip URLs and callbacks to create matching pairs
        for url, callback in zip(links.values(), callbacks):
            yield scrapy.Request(url, callback=callback)

Alternative: Return a List Comprehension

If you’d rather stick to returning a list (like your original code), a list comprehension works too—it’s equivalent to your manual list but avoids repetition:

class SiteFetching(scrapy.Spider):
    name = 'Site'
    
    def start_requests(self):
        links = {
            'transcription_page': 'https://www.rev.com/freelancers/transcription', 
            'captions_page': 'https://www.rev.com/freelancers/captions', 
            'subtitles_page': 'https://www.rev.com/freelancers/subtitles'
        }
        callbacks = [self.parse_transcription, self.parse_caption, self.parse_subtitles]
        
        return [scrapy.Request(url, callback=callback) for url, callback in zip(links.values(), callbacks)]

Why This Works

  • Using yield lets Scrapy iterate over each request as it’s generated, ensuring all three URLs get processed.
  • zip() guarantees each URL is matched with the correct callback, so you never mix up which parser handles which page.
  • All these approaches eliminate repetitive code, making it easier to add new URLs/callbacks later without rewriting the same lines.

内容的提问来源于stack exchange,提问作者Zaryab

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 07:22:42