You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Scrapy提取Zomato餐厅页面链接?爬取结果异常排查

Hey there, the issue with your current spider is that your CSS selector .col-s-12 is way too broad—it targets every element with that class, which includes tons of non-restaurant links (like pagination, navigation, ads, etc.). Let's tweak this to only grab the actual restaurant page URLs.

Here are two solid solutions you can use:

Solution 1: Use a Precise CSS Selector

First, inspect the Zomato search results page with your browser's dev tools (right-click a restaurant card → Inspect). You'll notice each restaurant is wrapped in a specific container with a unique class (right now on the Turkish Zomato site, restaurant cards use the class sc-1hp8d8a-0). We can target that container to only pull links from restaurant cards:

import scrapy

class ZomatoSpider(scrapy.Spider):
    name = 'zomato'
    allowed_domains = ["zomato.com"]
    start_urls = ['https://www.zomato.com/tr/istanbul/restoranlar?page=1']

    def parse(self, response):
        # Target the specific container holding each restaurant card
        restaurant_cards = response.css('.sc-1hp8d8a-0')
        
        for card in restaurant_cards:
            # Extract the relative link from the card's <a> tag
            relative_link = card.css('a::attr(href)').get()
            
            if relative_link:
                # Convert relative path to full, valid URL
                full_restaurant_link = response.urljoin(relative_link)
                yield {'restaurant_link': full_restaurant_link}

Note: If the class sc-1hp8d8a-0 changes (Zomato updates their frontend), re-inspect the page to get the current container class.

If you prefer a more resilient approach (in case Zomato changes their container classes), you can filter links based on the unique format of restaurant URLs. Restaurant links on Zomato follow this pattern: /istanbul/[restaurant-slug]-istanbul (no query parameters like ?page=).

We can use a regex to match this pattern directly:

import scrapy
import re

class ZomatoSpider(scrapy.Spider):
    name = 'zomato'
    allowed_domains = ["zomato.com"]
    start_urls = ['https://www.zomato.com/tr/istanbul/restoranlar?page=1']
    
    # Regex to match valid restaurant relative links
    RESTAURANT_LINK_PATTERN = re.compile(r'^/istanbul/[^?]+$')

    def parse(self, response):
        # Extract all links, then filter using our regex
        all_links = response.css('a::attr(href)').extract()
        
        for link in all_links:
            if self.RESTAURANT_LINK_PATTERN.match(link):
                full_link = response.urljoin(link)
                yield {'restaurant_link': full_link}

Bonus: Even Cleaner Regex Filtering with Scrapy's re() Method

Scrapy lets you combine CSS selection and regex in one step for efficiency:

import scrapy

class ZomatoSpider(scrapy.Spider):
    name = 'zomato'
    allowed_domains = ["zomato.com"]
    start_urls = ['https://www.zomato.com/tr/istanbul/restoranlar?page=1']

    def parse(self, response):
        # Directly extract links matching our regex pattern
        restaurant_links = response.css('a::attr(href)').re(r'^/istanbul/[^?]+$')
        
        for link in restaurant_links:
            full_link = response.urljoin(link)
            yield {'restaurant_link': full_link}

Any of these approaches will ensure you only get the restaurant page links you need, instead of every random link on the page.

内容的提问来源于stack exchange,提问作者cagahan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 11:32:38