Scrapy新手求助:爬虫未爬取到数据(爬取0页),疑XPath或URL GET参数问题
Hey there! Let's work through your Scrapy issues step by step to get you pulling those apartment titles successfully.
First: Fixing the GET Parameter Problem
You don't need to handle GET parameters separately like you would with POST data—just put the full URL (with all your filters) directly into start_urls. Right now, your spider is only visiting the default apartment search page (without your Prague, 2+kt, floor filters), which is why you're not seeing the data you want. Update your start_urls to:
start_urls = ["https://www.sreality.cz/en/search/for-sale/apartments/praha?disposition=2%2Bkt&published=month&min-floor=1&max-floor=3"]
Second: Why Your XPath Isn't Working (The Hidden Issue)
Your XPath has two critical problems:
- Typo:
//basci/h2/title/@contenthas a typo (basciinstead ofbasic), but even fixing that won't help. - Dynamic Content: This site uses AngularJS to load apartment data via AJAX. If you right-click the page and select "View Page Source", you'll notice the
name ng-bindingspan tags (and apartment titles) don't exist in the static HTML—they're generated after the browser runs the site's JavaScript. Scrapy's default behavior only fetches static HTML, so it can't see these dynamically rendered elements.
The Best Solution: Scrape the Site's API (Faster & More Reliable)
Instead of dealing with JavaScript rendering, you can scrape the direct API that the site uses to load its data. Here's how to set that up:
- Open your browser's DevTools (F12), go to the Network tab, and filter for XHR/Fetch requests.
- Refresh your target page, and you'll see a request to an API endpoint like
https://www.sreality.cz/api/en/v2/estateswith all your filter parameters included. - Use this API URL in your spider—it returns clean JSON with all the listing details, including titles.
Here's your revised spider using the API:
import scrapy import json class Sp1Spider(scrapy.Spider): name = 'sp1' allowed_domains = ['www.sreality.cz'] # API endpoint with your filters pre-applied start_urls = [ "https://www.sreality.cz/api/en/v2/estates?category_main_cb=1&category_type_cb=1&locality_region_id=10&disposition=2%2Bkt&published=month&min_floor=1&max_floor=3&per_page=60" ] def parse(self, response): # Parse the JSON response data = json.loads(response.text) # Loop through each estate in the response for estate in data['_embedded']['estates']: yield { 'apartment_title': estate['name'] } # Optional: Add pagination to scrape more results if 'next' in data['_links']: next_page_url = data['_links']['next']['href'] yield scrapy.Request(url=next_page_url, callback=self.parse)
If You Want to Scrape the Rendered Page (Less Efficient)
If you prefer to scrape the actual web page instead of the API, you'll need a tool to render JavaScript. Scrapy works with libraries like scrapy-playwright for this:
- Install it first:
pip install scrapy-playwright - Enable it in your
settings.py:DOWNLOAD_HANDLERS = { "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler", } PLAYWRIGHT_LAUNCH_OPTIONS = {"headless": True} - Update your spider to use Playwright for rendering:
import scrapy class Sp1Spider(scrapy.Spider): name = 'sp1' allowed_domains = ['www.sreality.cz'] start_urls = [ "https://www.sreality.cz/en/search/for-sale/apartments/praha?disposition=2%2Bkt&published=month&min-floor=1&max-floor=3" ] def start_requests(self): for url in self.start_urls: # Tell Scrapy to use Playwright for this request yield scrapy.Request(url, meta={"playwright": True}) def parse(self, response): # Now the response includes dynamically rendered content apartment_titles = response.xpath('//span[@class="name ng-binding"]/text()').getall() for title in apartment_titles: yield {"apartment_title": title.strip()}
Quick Recap
- You weren't accessing the correct filtered page (fixed by updating
start_urlswith your full URL) - Static scraping couldn't see the JS-rendered apartment titles (fixed by using the API or JS rendering tools)
- Your original XPath had a typo, which we've corrected in the examples above
内容的提问来源于stack exchange,提问作者Iyad

