You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取JSON加载无限滚动的17000条详情页数据求助

Solution: Scrapy Spider for 1177.se API Pagination

Got it, let's tackle this problem step by step. Since the site ignores the batchsize parameter and uses a JSON API with the p (page) parameter for pagination, we can build a Scrapy spider that directly targets these API endpoints, iterates through pages, and fetches the 17000 detail pages you need. Here's a fully runnable initial implementation:

import scrapy
from scrapy.http import JsonRequest

class HjvSpider(scrapy.Spider):
    name = 'hjv_spider'
    allowed_domains = ['1177.se']
    # Base API URL with placeholders for the page parameter
    base_api_url = 'https://www.1177.se/api/hjv/search?batchsize=10&caretype=&componentname&cs=false&location=&p={}&q=&s=name&sortorder=name&st=4af2ed43-1154-4363-ae6b-718f9b84d23a'
    total_items_needed = 17000
    items_collected = 0

    def start_requests(self):
        # Start with page 1
        yield JsonRequest(
            url=self.base_api_url.format(1),
            headers={
                'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
                'Accept': 'application/json, text/plain, */*'
            },
            callback=self.parse_api_response,
            meta={'page': 1}
        )

    def parse_api_response(self, response):
        data = response.json()
        results = data.get('Results', [])

        for result in results:
            # Adjust the 'Url' field to match the actual key in the JSON response
            detail_url = result.get('Url')
            if detail_url:
                # Make sure the URL is absolute (if it's relative)
                if not detail_url.startswith('http'):
                    detail_url = f'https://www.1177.se{detail_url}'
                yield scrapy.Request(
                    url=detail_url,
                    callback=self.parse_detail_page
                )
                self.items_collected += 1
                # Stop if we've collected enough items
                if self.items_collected >= self.total_items_needed:
                    return

        # Check if there's a next page to crawl
        next_page = data.get('NextPage')
        current_page = response.meta['page']
        if next_page and self.items_collected < self.total_items_needed:
            next_page_num = current_page + 1
            yield JsonRequest(
                url=self.base_api_url.format(next_page_num),
                headers={
                    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
                    'Accept': 'application/json, text/plain, */*'
                },
                callback=self.parse_api_response,
                meta={'page': next_page_num}
            )

    def parse_detail_page(self, response):
        # Extract data from the detail page here
        # Example: extract title, address, etc.
        yield {
            'title': response.css('h1::text').get(),
            'url': response.url,
            # Add other fields you need to scrape
        }

Key Details & Adjustments:

  • Headers: The User-Agent and Accept headers help mimic a browser request, which reduces the chance of being blocked by the API. You can update the User-Agent to match your current browser's if needed.
  • JSON Response Parsing: The spider assumes the API returns a Results array with your target items, and each item has a Url field pointing to the detail page. You'll need to check the actual JSON structure (use browser dev tools' Network tab) and adjust the field names if they differ (e.g., maybe DetailUrl instead of Url).
  • Pagination Control: We check the NextPage field from the API response to decide if we should crawl the next page. We also track the total number of items collected to stop once we hit 17000.
  • Detail Page Parsing: The parse_detail_page method is a placeholder—replace the example fields with the actual data you need to extract from the detail pages (using CSS or XPath selectors).

Tips for Avoiding Blocks:

  • Add a DOWNLOAD_DELAY = 1 (or higher) in your settings.py to slow down requests and avoid overwhelming the server.
  • If you encounter rate limiting, consider adding rotating proxies or using Scrapy's AutoThrottle extension.

内容的提问来源于stack exchange,提问作者Adam Robinsson

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 15:02:47