You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何为输出指定格式的自定义Scrapy爬虫编写测试用例

How to Write Test Cases for Your Scrapy Spider

Absolutely! You can absolutely write test cases for your Scrapy spider—Scrapy plays nicely with Python's standard unittest framework, and you can mock website responses to avoid hitting the real Adapt.io site during tests (which is faster and avoids triggering any anti-scraping measures). Let's walk through two practical approaches:

This focuses on testing your spider's core logic without making real HTTP requests. We'll create a mock HTML response that matches the structure of the page your spider targets, then pass it to the parse method and verify the output items are correct.

Step 1: Create the Test File

Create a file like test_spider.py in your Scrapy project directory. Here's the full test code:

import unittest
import scrapy
from scrapy.http import HtmlResponse
from your_spider_file import CompanySpider  # Replace with your actual spider filename (e.g., spiders/company.py)

class TestCompanySpider(unittest.TestCase):
    def setUp(self):
        # Initialize your spider once before each test
        self.spider = CompanySpider()

    def test_parse_output(self):
        # Mock HTML that mirrors the structure of the Adapt.io directory page
        mock_page_html = """
        <div class="DirectoryList_link">
            <a href="https://www.adapt.io/company/a--communications-and-security">A + Communications and Security</a>
        </div>
        <div class="DirectoryList_link">
            <a href="https://www.adapt.io/company/a-a-technology-group">A&A Technology Group</a>
        </div>
        """
        # Create a mock Scrapy response object
        mock_response = HtmlResponse(
            url=self.spider.start_urls[0],
            body=mock_page_html.encode('utf-8'),
            encoding='utf-8'
        )

        # Run the parse method and collect the yielded items
        parsed_items = list(self.spider.parse(mock_response))

        # Verify the number of items matches our mock data
        self.assertEqual(len(parsed_items), 2)

        # Check the first item's data is correct
        first_company = parsed_items[0]
        self.assertEqual(first_company['company_name'], 'A + Communications and Security')
        self.assertEqual(first_company['source_url'], '/company/a--communications-and-security')

        # Check the second item's data is correct
        second_company = parsed_items[1]
        self.assertEqual(second_company['company_name'], 'A&A Technology Group')
        self.assertEqual(second_company['source_url'], '/company/a-a-technology-group')

if __name__ == '__main__':
    unittest.main()

How This Works:

  • We use HtmlResponse to simulate the page your spider would receive.
  • We convert the generator returned by parse into a list to inspect the items.
  • We use unittest assertions to confirm the output matches your expected format and values.

2. Integration Testing (Optional)

If you want to test the full pipeline (including making real HTTP requests), you can use Scrapy's CrawlerProcess to run the spider and capture its output. Note: Only run this occasionally, as it hits the real website.

from scrapy.crawler import CrawlerProcess
from your_spider_file import CompanySpider

def test_full_spider_run():
    collected_items = []

    # Define a callback to collect items as they're yielded
    def item_collector(item):
        collected_items.append(item)

    # Set up the crawler with your spider's settings
    process = CrawlerProcess(settings={
        'TELNETCONSOLE_ENABLED': False,
        'ROBOTSTXT_OBEY': False,
        'LOG_LEVEL': 'ERROR'  # Reduce log noise during testing
    })

    # Start the crawl and block until it finishes
    process.crawl(CompanySpider, item_pipeline=item_collector)
    process.start()

    # Basic assertions (adjust based on expected results)
    assert len(collected_items) > 0, "No items were scraped"
    assert all('company_name' in item and 'source_url' in item for item in collected_items), "Missing fields in items"

if __name__ == '__main__':
    test_full_spider_run()

Key Tips for Testing

  • Keep Mock HTML Accurate: If Adapt.io changes their page structure, update your mock HTML to match—otherwise your tests will fail even if the spider works.
  • Test Edge Cases: Add mock entries for company names with special characters, or URLs that might break your split logic, to ensure your spider handles them correctly.
  • Separate Unit vs Integration Tests: Use unit tests for daily development (fast, no network calls) and integration tests only to validate against the real site periodically.

内容的提问来源于stack exchange,提问作者grroom model

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 07:38:11