如何为输出指定格式的自定义Scrapy爬虫编写测试用例
Absolutely! You can absolutely write test cases for your Scrapy spider—Scrapy plays nicely with Python's standard unittest framework, and you can mock website responses to avoid hitting the real Adapt.io site during tests (which is faster and avoids triggering any anti-scraping measures). Let's walk through two practical approaches:
1. Unit Testing the parse Method (Recommended)
This focuses on testing your spider's core logic without making real HTTP requests. We'll create a mock HTML response that matches the structure of the page your spider targets, then pass it to the parse method and verify the output items are correct.
Step 1: Create the Test File
Create a file like test_spider.py in your Scrapy project directory. Here's the full test code:
import unittest import scrapy from scrapy.http import HtmlResponse from your_spider_file import CompanySpider # Replace with your actual spider filename (e.g., spiders/company.py) class TestCompanySpider(unittest.TestCase): def setUp(self): # Initialize your spider once before each test self.spider = CompanySpider() def test_parse_output(self): # Mock HTML that mirrors the structure of the Adapt.io directory page mock_page_html = """ <div class="DirectoryList_link"> <a href="https://www.adapt.io/company/a--communications-and-security">A + Communications and Security</a> </div> <div class="DirectoryList_link"> <a href="https://www.adapt.io/company/a-a-technology-group">A&A Technology Group</a> </div> """ # Create a mock Scrapy response object mock_response = HtmlResponse( url=self.spider.start_urls[0], body=mock_page_html.encode('utf-8'), encoding='utf-8' ) # Run the parse method and collect the yielded items parsed_items = list(self.spider.parse(mock_response)) # Verify the number of items matches our mock data self.assertEqual(len(parsed_items), 2) # Check the first item's data is correct first_company = parsed_items[0] self.assertEqual(first_company['company_name'], 'A + Communications and Security') self.assertEqual(first_company['source_url'], '/company/a--communications-and-security') # Check the second item's data is correct second_company = parsed_items[1] self.assertEqual(second_company['company_name'], 'A&A Technology Group') self.assertEqual(second_company['source_url'], '/company/a-a-technology-group') if __name__ == '__main__': unittest.main()
How This Works:
- We use
HtmlResponseto simulate the page your spider would receive. - We convert the generator returned by
parseinto a list to inspect the items. - We use
unittestassertions to confirm the output matches your expected format and values.
2. Integration Testing (Optional)
If you want to test the full pipeline (including making real HTTP requests), you can use Scrapy's CrawlerProcess to run the spider and capture its output. Note: Only run this occasionally, as it hits the real website.
from scrapy.crawler import CrawlerProcess from your_spider_file import CompanySpider def test_full_spider_run(): collected_items = [] # Define a callback to collect items as they're yielded def item_collector(item): collected_items.append(item) # Set up the crawler with your spider's settings process = CrawlerProcess(settings={ 'TELNETCONSOLE_ENABLED': False, 'ROBOTSTXT_OBEY': False, 'LOG_LEVEL': 'ERROR' # Reduce log noise during testing }) # Start the crawl and block until it finishes process.crawl(CompanySpider, item_pipeline=item_collector) process.start() # Basic assertions (adjust based on expected results) assert len(collected_items) > 0, "No items were scraped" assert all('company_name' in item and 'source_url' in item for item in collected_items), "Missing fields in items" if __name__ == '__main__': test_full_spider_run()
Key Tips for Testing
- Keep Mock HTML Accurate: If Adapt.io changes their page structure, update your mock HTML to match—otherwise your tests will fail even if the spider works.
- Test Edge Cases: Add mock entries for company names with special characters, or URLs that might break your
splitlogic, to ensure your spider handles them correctly. - Separate Unit vs Integration Tests: Use unit tests for daily development (fast, no network calls) and integration tests only to validate against the real site periodically.
内容的提问来源于stack exchange,提问作者grroom model

