You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy新手求助:爬虫未爬取到数据(爬取0页),疑XPath或URL GET参数问题

Hey there! Let's work through your Scrapy issues step by step to get you pulling those apartment titles successfully.

First: Fixing the GET Parameter Problem

You don't need to handle GET parameters separately like you would with POST data—just put the full URL (with all your filters) directly into start_urls. Right now, your spider is only visiting the default apartment search page (without your Prague, 2+kt, floor filters), which is why you're not seeing the data you want. Update your start_urls to:

start_urls = ["https://www.sreality.cz/en/search/for-sale/apartments/praha?disposition=2%2Bkt&published=month&min-floor=1&max-floor=3"]

Second: Why Your XPath Isn't Working (The Hidden Issue)

Your XPath has two critical problems:

  1. Typo: //basci/h2/title/@content has a typo (basci instead of basic), but even fixing that won't help.
  2. Dynamic Content: This site uses AngularJS to load apartment data via AJAX. If you right-click the page and select "View Page Source", you'll notice the name ng-binding span tags (and apartment titles) don't exist in the static HTML—they're generated after the browser runs the site's JavaScript. Scrapy's default behavior only fetches static HTML, so it can't see these dynamically rendered elements.

The Best Solution: Scrape the Site's API (Faster & More Reliable)

Instead of dealing with JavaScript rendering, you can scrape the direct API that the site uses to load its data. Here's how to set that up:

  1. Open your browser's DevTools (F12), go to the Network tab, and filter for XHR/Fetch requests.
  2. Refresh your target page, and you'll see a request to an API endpoint like https://www.sreality.cz/api/en/v2/estates with all your filter parameters included.
  3. Use this API URL in your spider—it returns clean JSON with all the listing details, including titles.

Here's your revised spider using the API:

import scrapy
import json

class Sp1Spider(scrapy.Spider):
    name = 'sp1'
    allowed_domains = ['www.sreality.cz']
    # API endpoint with your filters pre-applied
    start_urls = [
        "https://www.sreality.cz/api/en/v2/estates?category_main_cb=1&category_type_cb=1&locality_region_id=10&disposition=2%2Bkt&published=month&min_floor=1&max_floor=3&per_page=60"
    ]

    def parse(self, response):
        # Parse the JSON response
        data = json.loads(response.text)
        # Loop through each estate in the response
        for estate in data['_embedded']['estates']:
            yield {
                'apartment_title': estate['name']
            }
        # Optional: Add pagination to scrape more results
        if 'next' in data['_links']:
            next_page_url = data['_links']['next']['href']
            yield scrapy.Request(url=next_page_url, callback=self.parse)

If You Want to Scrape the Rendered Page (Less Efficient)

If you prefer to scrape the actual web page instead of the API, you'll need a tool to render JavaScript. Scrapy works with libraries like scrapy-playwright for this:

  1. Install it first: pip install scrapy-playwright
  2. Enable it in your settings.py:
    DOWNLOAD_HANDLERS = {
        "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
        "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    }
    PLAYWRIGHT_LAUNCH_OPTIONS = {"headless": True}
    
  3. Update your spider to use Playwright for rendering:
    import scrapy
    
    class Sp1Spider(scrapy.Spider):
        name = 'sp1'
        allowed_domains = ['www.sreality.cz']
        start_urls = [
            "https://www.sreality.cz/en/search/for-sale/apartments/praha?disposition=2%2Bkt&published=month&min-floor=1&max-floor=3"
        ]
    
        def start_requests(self):
            for url in self.start_urls:
                # Tell Scrapy to use Playwright for this request
                yield scrapy.Request(url, meta={"playwright": True})
    
        def parse(self, response):
            # Now the response includes dynamically rendered content
            apartment_titles = response.xpath('//span[@class="name ng-binding"]/text()').getall()
            for title in apartment_titles:
                yield {"apartment_title": title.strip()}
    

Quick Recap

  • You weren't accessing the correct filtered page (fixed by updating start_urls with your full URL)
  • Static scraping couldn't see the JS-rendered apartment titles (fixed by using the API or JS rendering tools)
  • Your original XPath had a typo, which we've corrected in the examples above

内容的提问来源于stack exchange,提问作者Iyad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.27 12:54:09