如何用Scrapy提取Zomato餐厅页面链接?爬取结果异常排查
Hey there, the issue with your current spider is that your CSS selector .col-s-12 is way too broad—it targets every element with that class, which includes tons of non-restaurant links (like pagination, navigation, ads, etc.). Let's tweak this to only grab the actual restaurant page URLs.
Here are two solid solutions you can use:
Solution 1: Use a Precise CSS Selector
First, inspect the Zomato search results page with your browser's dev tools (right-click a restaurant card → Inspect). You'll notice each restaurant is wrapped in a specific container with a unique class (right now on the Turkish Zomato site, restaurant cards use the class sc-1hp8d8a-0). We can target that container to only pull links from restaurant cards:
import scrapy class ZomatoSpider(scrapy.Spider): name = 'zomato' allowed_domains = ["zomato.com"] start_urls = ['https://www.zomato.com/tr/istanbul/restoranlar?page=1'] def parse(self, response): # Target the specific container holding each restaurant card restaurant_cards = response.css('.sc-1hp8d8a-0') for card in restaurant_cards: # Extract the relative link from the card's <a> tag relative_link = card.css('a::attr(href)').get() if relative_link: # Convert relative path to full, valid URL full_restaurant_link = response.urljoin(relative_link) yield {'restaurant_link': full_restaurant_link}
Note: If the class sc-1hp8d8a-0 changes (Zomato updates their frontend), re-inspect the page to get the current container class.
Solution 2: Filter Links with Regular Expressions
If you prefer a more resilient approach (in case Zomato changes their container classes), you can filter links based on the unique format of restaurant URLs. Restaurant links on Zomato follow this pattern: /istanbul/[restaurant-slug]-istanbul (no query parameters like ?page=).
We can use a regex to match this pattern directly:
import scrapy import re class ZomatoSpider(scrapy.Spider): name = 'zomato' allowed_domains = ["zomato.com"] start_urls = ['https://www.zomato.com/tr/istanbul/restoranlar?page=1'] # Regex to match valid restaurant relative links RESTAURANT_LINK_PATTERN = re.compile(r'^/istanbul/[^?]+$') def parse(self, response): # Extract all links, then filter using our regex all_links = response.css('a::attr(href)').extract() for link in all_links: if self.RESTAURANT_LINK_PATTERN.match(link): full_link = response.urljoin(link) yield {'restaurant_link': full_link}
Bonus: Even Cleaner Regex Filtering with Scrapy's re() Method
Scrapy lets you combine CSS selection and regex in one step for efficiency:
import scrapy class ZomatoSpider(scrapy.Spider): name = 'zomato' allowed_domains = ["zomato.com"] start_urls = ['https://www.zomato.com/tr/istanbul/restoranlar?page=1'] def parse(self, response): # Directly extract links matching our regex pattern restaurant_links = response.css('a::attr(href)').re(r'^/istanbul/[^?]+$') for link in restaurant_links: full_link = response.urljoin(link) yield {'restaurant_link': full_link}
Any of these approaches will ensure you only get the restaurant page links you need, instead of every random link on the page.
内容的提问来源于stack exchange,提问作者cagahan

