如何在Scrapy的start_requests中用循环返回多个scrapy.Request避免重复代码?
start_requests Method Hey there! I get it—repeating scrapy.Request lines is tedious, and it makes total sense to use a loop to generate all your requests cleanly. Let's break down why your initial loop attempt might have failed, and get you set up with scalable, DRY (Don't Repeat Yourself) code.
The Core Issue
Your current code works, but manually listing three scrapy.Request objects isn't ideal. If you tried a loop before and only got the first result, chances are you either returned a single request instead of an iterable, or your loop didn't properly pair URLs with their matching callback functions.
Solution 1: Use a Generator with yield (Recommended)
Scrapy’s start_requests method can act as a generator—using yield for each request aligns perfectly with how Scrapy handles asynchronous tasks, and it’s super readable. Here’s how to rewrite your code:
class SiteFetching(scrapy.Spider): name = 'Site' def start_requests(self): # Pair each URL directly with its corresponding callback url_callback_pairs = [ ('https://www.rev.com/freelancers/transcription', self.parse_transcription), ('https://www.rev.com/freelancers/captions', self.parse_caption), ('https://www.rev.com/freelancers/subtitles', self.parse_subtitles) ] # Loop through each pair and yield the request for url, callback in url_callback_pairs: yield scrapy.Request(url, callback=callback)
Solution 2: Keep Your Original Dictionary (If You Prefer)
If you want to retain your links dictionary, just use zip() to pair its values with your callback list correctly:
class SiteFetching(scrapy.Spider): name = 'Site' def start_requests(self): links = { 'transcription_page': 'https://www.rev.com/freelancers/transcription', 'captions_page': 'https://www.rev.com/freelancers/captions', 'subtitles_page': 'https://www.rev.com/freelancers/subtitles' } callbacks = [self.parse_transcription, self.parse_caption, self.parse_subtitles] # Zip URLs and callbacks to create matching pairs for url, callback in zip(links.values(), callbacks): yield scrapy.Request(url, callback=callback)
Alternative: Return a List Comprehension
If you’d rather stick to returning a list (like your original code), a list comprehension works too—it’s equivalent to your manual list but avoids repetition:
class SiteFetching(scrapy.Spider): name = 'Site' def start_requests(self): links = { 'transcription_page': 'https://www.rev.com/freelancers/transcription', 'captions_page': 'https://www.rev.com/freelancers/captions', 'subtitles_page': 'https://www.rev.com/freelancers/subtitles' } callbacks = [self.parse_transcription, self.parse_caption, self.parse_subtitles] return [scrapy.Request(url, callback=callback) for url, callback in zip(links.values(), callbacks)]
Why This Works
- Using
yieldlets Scrapy iterate over each request as it’s generated, ensuring all three URLs get processed. zip()guarantees each URL is matched with the correct callback, so you never mix up which parser handles which page.- All these approaches eliminate repetitive code, making it easier to add new URLs/callbacks later without rewriting the same lines.
内容的提问来源于stack exchange,提问作者Zaryab

