Scrapy单爬虫多域名爬取:从数据库加载URL及解析适配问询
Hey there! Since you're new to Python and Scrapy, let's break down your two key problems with practical, easy-to-follow solutions:
1. Loading URLs to Crawl from a Database
Scrapy's default start_urls list works great for static URLs, but for dynamic sources like databases, you'll want to override the start_requests() method. Here's how to do it smoothly:
First, store your database credentials in Scrapy's settings.py to keep your code clean and configurable:
# settings.py DATABASE = { 'drivername': 'postgresql', # Swap for 'sqlite' or 'mysql' based on your DB 'host': 'localhost', 'port': '5432', 'username': 'your_db_user', 'password': 'your_db_pass', 'database': 'credit_card_db' }
Then, in your spider, connect to the database, fetch the URLs (plus extra metadata like bank name for later routing), and yield Request objects for each entry:
# your_spider.py import scrapy import psycopg2 from scrapy.utils.project import get_project_settings class CreditCardSpider(scrapy.Spider): name = 'credit_card_spider' def start_requests(self): settings = get_project_settings() db_params = settings.get('DATABASE') # Connect to your database conn = psycopg2.connect( host=db_params['host'], port=db_params['port'], user=db_params['username'], password=db_params['password'], dbname=db_params['database'] ) cursor = conn.cursor() # Query URLs and associated bank names (critical for routing later!) cursor.execute("SELECT url, bank_name FROM credit_card_pages") rows = cursor.fetchall() for url, bank_name in rows: # Pass bank name in meta to use for parsing routing yield scrapy.Request( url=url, meta={'bank_name': bank_name}, callback=self.parse_router ) # Clean up database connections cursor.close() conn.close()
2. Routing to Specific Parse Methods When Site Structures Change
Each bank's site has its own unique HTML structure, so you need a way to send each URL to a tailored parse function. Here are two reliable approaches:
Option 1: Use a Router Function with Meta Data
Define a parse_router method that checks the bank name from your database and routes the response to the correct parse function:
# Add these methods to your CreditCardSpider class def parse_router(self, response): bank_name = response.meta.get('bank_name') # Map bank names to their dedicated parse functions parse_map = { 'chase': self.parse_chase_card, 'citi': self.parse_citi_card, 'amex': self.parse_amex_card, # Add more banks as you expand your crawler } # Fallback to a generic parser if no specific one exists parse_func = parse_map.get(bank_name.lower(), self.parse_generic_card) yield from parse_func(response) def parse_chase_card(self, response): # Chase-specific selectors to extract card data yield { 'card_name': response.css('h1.card-product-title::text').get().strip(), 'annual_fee': response.css('.annual-fee-value::text').get().strip(), 'apr_rate': response.css('.variable-apr-rate::text').get().strip(), 'card_image': response.css('.card-product-image img::attr(src)').get() } def parse_citi_card(self, response): # Citi-specific selectors yield { 'card_name': response.css('.card-header h2::text').get().strip(), 'annual_fee': response.css('.fee-details span::text').get().strip(), 'apr_rate': response.css('.apr-info .rate-number::text').get().strip(), 'card_image': response.css('.card-visual img::attr(data-src)').get() } def parse_generic_card(self, response): # Fallback for unconfigured banks self.logger.warning(f"No specific parser found for {response.meta.get('bank_name')}") yield { 'card_name': response.css('title::text').get(), 'url': response.url, # Leave other fields empty or add generic selectors as needed }
Option 2: Assign Callback Directly in start_requests
If you prefer a more direct approach, skip the router and assign the callback function when yielding the Request:
# Inside the start_requests loop parse_map = { 'chase': self.parse_chase_card, 'citi': self.parse_citi_card, 'amex': self.parse_amex_card, } callback = parse_map.get(bank_name.lower(), self.parse_generic_card) yield scrapy.Request( url=url, callback=callback )
This way, each URL goes straight to its dedicated parse function without a middle step. Both methods let you easily update or add parsers when a bank changes its site structure—no need to rewrite the entire spider.
内容的提问来源于stack exchange,提问作者Vishal Patel

