Scrapy爬取JSON加载无限滚动的17000条详情页数据求助
Solution: Scrapy Spider for 1177.se API Pagination
Got it, let's tackle this problem step by step. Since the site ignores the batchsize parameter and uses a JSON API with the p (page) parameter for pagination, we can build a Scrapy spider that directly targets these API endpoints, iterates through pages, and fetches the 17000 detail pages you need. Here's a fully runnable initial implementation:
import scrapy from scrapy.http import JsonRequest class HjvSpider(scrapy.Spider): name = 'hjv_spider' allowed_domains = ['1177.se'] # Base API URL with placeholders for the page parameter base_api_url = 'https://www.1177.se/api/hjv/search?batchsize=10&caretype=&componentname&cs=false&location=&p={}&q=&s=name&sortorder=name&st=4af2ed43-1154-4363-ae6b-718f9b84d23a' total_items_needed = 17000 items_collected = 0 def start_requests(self): # Start with page 1 yield JsonRequest( url=self.base_api_url.format(1), headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'application/json, text/plain, */*' }, callback=self.parse_api_response, meta={'page': 1} ) def parse_api_response(self, response): data = response.json() results = data.get('Results', []) for result in results: # Adjust the 'Url' field to match the actual key in the JSON response detail_url = result.get('Url') if detail_url: # Make sure the URL is absolute (if it's relative) if not detail_url.startswith('http'): detail_url = f'https://www.1177.se{detail_url}' yield scrapy.Request( url=detail_url, callback=self.parse_detail_page ) self.items_collected += 1 # Stop if we've collected enough items if self.items_collected >= self.total_items_needed: return # Check if there's a next page to crawl next_page = data.get('NextPage') current_page = response.meta['page'] if next_page and self.items_collected < self.total_items_needed: next_page_num = current_page + 1 yield JsonRequest( url=self.base_api_url.format(next_page_num), headers={ 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'application/json, text/plain, */*' }, callback=self.parse_api_response, meta={'page': next_page_num} ) def parse_detail_page(self, response): # Extract data from the detail page here # Example: extract title, address, etc. yield { 'title': response.css('h1::text').get(), 'url': response.url, # Add other fields you need to scrape }
Key Details & Adjustments:
- Headers: The
User-AgentandAcceptheaders help mimic a browser request, which reduces the chance of being blocked by the API. You can update the User-Agent to match your current browser's if needed. - JSON Response Parsing: The spider assumes the API returns a
Resultsarray with your target items, and each item has aUrlfield pointing to the detail page. You'll need to check the actual JSON structure (use browser dev tools' Network tab) and adjust the field names if they differ (e.g., maybeDetailUrlinstead ofUrl). - Pagination Control: We check the
NextPagefield from the API response to decide if we should crawl the next page. We also track the total number of items collected to stop once we hit 17000. - Detail Page Parsing: The
parse_detail_pagemethod is a placeholder—replace the example fields with the actual data you need to extract from the detail pages (using CSS or XPath selectors).
Tips for Avoiding Blocks:
- Add a
DOWNLOAD_DELAY = 1(or higher) in yoursettings.pyto slow down requests and avoid overwhelming the server. - If you encounter rate limiting, consider adding rotating proxies or using Scrapy's
AutoThrottleextension.
内容的提问来源于stack exchange,提问作者Adam Robinsson
相关产品推荐
相关产品推荐

