You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy爬取JSON网页?StarCityGames爬取无数据求助

Let's break down why your Scrapy spider isn't returning any data and fix it step by step:

1. Fix the allowed_domains configuration

Your current allowed_domains is set to the full URL, which is incorrect. Scrapy expects only the domain name here—using the full URL will cause your requests to be filtered out by the framework:

allowed_domains = ['starcitygames.com']

2. Identify the correct data source

The URL you're targeting (http://www.starcitygames.com/buylist/search?search-type=category&id=5061) returns an HTML page, not raw JSON. Trying to parse this with json.loads() should throw an error, but if it's not, chances are the framework is swallowing exceptions, or you're misunderstanding where the data lives.

This page loads its product data dynamically via AJAX. To find the actual JSON API:

  • Open your browser's DevTools (F12) → go to the Network tab
  • Refresh the page and filter for XHR requests
  • Look for a request that returns the product list (it might look something like /buylist/search/data?search-type=category&id=5061)

3. Adjust your parsing logic to match the JSON structure

Once you have the correct JSON endpoint, update your start_urls to point to it. Then, check the structure of the JSON response—most APIs wrap data in a top-level field like items or data, not return a raw array. For example, if the response looks like this:

{
  "success": true,
  "items": [
    {"name": "Card Name", "condition": "Near Mint", "price": "0.50", "rarity": "Common"}
  ]
}

Update your parse method to target the right field:

def parse(self, response):
    # Use response.text instead of body_as_unicode() (it's the modern equivalent)
    jsonresponse = json.loads(response.text)
    # Iterate over the actual items array in the response
    for item_data in jsonresponse.get('items', []):
        loader = ItemLoader(item=NameItem())
        loader.default_input_processor = MapCompose(str)
        loader.default_output_processor = Join(' ')
        for (field, path) in self.jmes_paths.items():
            loader.add_value(field, SelectJmes(path)(item_data))
        yield loader.load_item()

4. Use Splash correctly (if needed)

You imported SplashRequest but aren't using it. If the JSON endpoint requires JavaScript rendering to be accessible, replace your start_urls with a start_requests method to use Splash:

def start_requests(self):
    # Replace with your actual JSON API URL
    url = 'http://www.starcitygames.com/buylist/search/data?search-type=category&id=5061'
    yield SplashRequest(url, self.parse, args={'wait': 0.5})

Also, make sure your settings.py has the required Splash configuration:

SPLASH_URL = 'http://localhost:8050'
DOWNLOADER_MIDDLEWARES = {
    'scrapy_splash.SplashCookiesMiddleware': 723,
    'scrapy_splash.SplashMiddleware': 725,
    'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
}
SPIDER_MIDDLEWARES = {
    'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
}
DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'

5. Debugging tips

  • Print the raw response first to confirm what you're getting: add print(response.text) at the start of your parse method.
  • Use the Scrapy shell to test requests interactively:
    scrapy shell "http://www.starcitygames.com/buylist/search?search-type=category&id=5061"
    
    Then run response.text to see the content, and json.loads(response.text) to test if parsing works.

内容的提问来源于stack exchange,提问作者Tom

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.13 09:00:10