You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy单爬虫多域名爬取:从数据库加载URL及解析适配问询

Scrapy Solutions for Your Multi-Bank Credit Card Crawler

Hey there! Since you're new to Python and Scrapy, let's break down your two key problems with practical, easy-to-follow solutions:

1. Loading URLs to Crawl from a Database

Scrapy's default start_urls list works great for static URLs, but for dynamic sources like databases, you'll want to override the start_requests() method. Here's how to do it smoothly:

First, store your database credentials in Scrapy's settings.py to keep your code clean and configurable:

# settings.py
DATABASE = {
    'drivername': 'postgresql', # Swap for 'sqlite' or 'mysql' based on your DB
    'host': 'localhost',
    'port': '5432',
    'username': 'your_db_user',
    'password': 'your_db_pass',
    'database': 'credit_card_db'
}

Then, in your spider, connect to the database, fetch the URLs (plus extra metadata like bank name for later routing), and yield Request objects for each entry:

# your_spider.py
import scrapy
import psycopg2
from scrapy.utils.project import get_project_settings

class CreditCardSpider(scrapy.Spider):
    name = 'credit_card_spider'

    def start_requests(self):
        settings = get_project_settings()
        db_params = settings.get('DATABASE')
        
        # Connect to your database
        conn = psycopg2.connect(
            host=db_params['host'],
            port=db_params['port'],
            user=db_params['username'],
            password=db_params['password'],
            dbname=db_params['database']
        )
        cursor = conn.cursor()
        
        # Query URLs and associated bank names (critical for routing later!)
        cursor.execute("SELECT url, bank_name FROM credit_card_pages")
        rows = cursor.fetchall()
        
        for url, bank_name in rows:
            # Pass bank name in meta to use for parsing routing
            yield scrapy.Request(
                url=url,
                meta={'bank_name': bank_name},
                callback=self.parse_router
            )
        
        # Clean up database connections
        cursor.close()
        conn.close()

2. Routing to Specific Parse Methods When Site Structures Change

Each bank's site has its own unique HTML structure, so you need a way to send each URL to a tailored parse function. Here are two reliable approaches:

Option 1: Use a Router Function with Meta Data

Define a parse_router method that checks the bank name from your database and routes the response to the correct parse function:

# Add these methods to your CreditCardSpider class
def parse_router(self, response):
    bank_name = response.meta.get('bank_name')
    
    # Map bank names to their dedicated parse functions
    parse_map = {
        'chase': self.parse_chase_card,
        'citi': self.parse_citi_card,
        'amex': self.parse_amex_card,
        # Add more banks as you expand your crawler
    }
    
    # Fallback to a generic parser if no specific one exists
    parse_func = parse_map.get(bank_name.lower(), self.parse_generic_card)
    yield from parse_func(response)

def parse_chase_card(self, response):
    # Chase-specific selectors to extract card data
    yield {
        'card_name': response.css('h1.card-product-title::text').get().strip(),
        'annual_fee': response.css('.annual-fee-value::text').get().strip(),
        'apr_rate': response.css('.variable-apr-rate::text').get().strip(),
        'card_image': response.css('.card-product-image img::attr(src)').get()
    }

def parse_citi_card(self, response):
    # Citi-specific selectors
    yield {
        'card_name': response.css('.card-header h2::text').get().strip(),
        'annual_fee': response.css('.fee-details span::text').get().strip(),
        'apr_rate': response.css('.apr-info .rate-number::text').get().strip(),
        'card_image': response.css('.card-visual img::attr(data-src)').get()
    }

def parse_generic_card(self, response):
    # Fallback for unconfigured banks
    self.logger.warning(f"No specific parser found for {response.meta.get('bank_name')}")
    yield {
        'card_name': response.css('title::text').get(),
        'url': response.url,
        # Leave other fields empty or add generic selectors as needed
    }

Option 2: Assign Callback Directly in start_requests

If you prefer a more direct approach, skip the router and assign the callback function when yielding the Request:

# Inside the start_requests loop
parse_map = {
    'chase': self.parse_chase_card,
    'citi': self.parse_citi_card,
    'amex': self.parse_amex_card,
}
callback = parse_map.get(bank_name.lower(), self.parse_generic_card)

yield scrapy.Request(
    url=url,
    callback=callback
)

This way, each URL goes straight to its dedicated parse function without a middle step. Both methods let you easily update or add parsers when a bank changes its site structure—no need to rewrite the entire spider.

内容的提问来源于stack exchange,提问作者Vishal Patel

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:38:28