如何用Scrapy爬取JSON网页?StarCityGames爬取无数据求助
Let's break down why your Scrapy spider isn't returning any data and fix it step by step:
1. Fix the allowed_domains configuration
Your current allowed_domains is set to the full URL, which is incorrect. Scrapy expects only the domain name here—using the full URL will cause your requests to be filtered out by the framework:
allowed_domains = ['starcitygames.com']
2. Identify the correct data source
The URL you're targeting (http://www.starcitygames.com/buylist/search?search-type=category&id=5061) returns an HTML page, not raw JSON. Trying to parse this with json.loads() should throw an error, but if it's not, chances are the framework is swallowing exceptions, or you're misunderstanding where the data lives.
This page loads its product data dynamically via AJAX. To find the actual JSON API:
- Open your browser's DevTools (F12) → go to the Network tab
- Refresh the page and filter for XHR requests
- Look for a request that returns the product list (it might look something like
/buylist/search/data?search-type=category&id=5061)
3. Adjust your parsing logic to match the JSON structure
Once you have the correct JSON endpoint, update your start_urls to point to it. Then, check the structure of the JSON response—most APIs wrap data in a top-level field like items or data, not return a raw array. For example, if the response looks like this:
{ "success": true, "items": [ {"name": "Card Name", "condition": "Near Mint", "price": "0.50", "rarity": "Common"} ] }
Update your parse method to target the right field:
def parse(self, response): # Use response.text instead of body_as_unicode() (it's the modern equivalent) jsonresponse = json.loads(response.text) # Iterate over the actual items array in the response for item_data in jsonresponse.get('items', []): loader = ItemLoader(item=NameItem()) loader.default_input_processor = MapCompose(str) loader.default_output_processor = Join(' ') for (field, path) in self.jmes_paths.items(): loader.add_value(field, SelectJmes(path)(item_data)) yield loader.load_item()
4. Use Splash correctly (if needed)
You imported SplashRequest but aren't using it. If the JSON endpoint requires JavaScript rendering to be accessible, replace your start_urls with a start_requests method to use Splash:
def start_requests(self): # Replace with your actual JSON API URL url = 'http://www.starcitygames.com/buylist/search/data?search-type=category&id=5061' yield SplashRequest(url, self.parse, args={'wait': 0.5})
Also, make sure your settings.py has the required Splash configuration:
SPLASH_URL = 'http://localhost:8050' DOWNLOADER_MIDDLEWARES = { 'scrapy_splash.SplashCookiesMiddleware': 723, 'scrapy_splash.SplashMiddleware': 725, 'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810, } SPIDER_MIDDLEWARES = { 'scrapy_splash.SplashDeduplicateArgsMiddleware': 100, } DUPEFILTER_CLASS = 'scrapy_splash.SplashAwareDupeFilter'
5. Debugging tips
- Print the raw response first to confirm what you're getting: add
print(response.text)at the start of yourparsemethod. - Use the Scrapy shell to test requests interactively:
Then runscrapy shell "http://www.starcitygames.com/buylist/search?search-type=category&id=5061"response.textto see the content, andjson.loads(response.text)to test if parsing works.
内容的提问来源于stack exchange,提问作者Tom

