You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬取zaubee.com:解决餐厅链接href为#的详情提取问题

问题描述

我正在开展一个基于Scrapy的网页爬取项目,目标是从zaubee.com网站采集餐厅经营信息。但页面中每个餐厅链接的href属性值均为“#”,导致无法跳转至详情页获取所需数据。我已编写初始Scrapy爬虫代码,运行时无法提取详情页的餐厅名称、网址、地址等信息,输出的详情字段均为None。现寻求可行解决方法,以实现访问餐厅详情页并采集名称、地址、电话、营业时间等信息的需求。

初始代码
import scrapy
from scrapy.linkextractors import LinkExtractor
from scrapy.spiders import CrawlSpider, Rule


class zaubeeSpider(scrapy.Spider):
    name = 'zaubeeerestaurant'
    allowed_domains = ['www.zaubee.com']
    start_urls = ['https://zaubee.com/category/restaurant-in-fredonia-hclq6jom']

def parse(self, response):
    restaurantlink = response.xpath("//div[@class='search-result__title-wrapper']/h2")
    for restaurant in restaurantlink:
        name= restaurant.xpath(".//text()").get()
        link = restaurant.xpath(".//@href").get()
        yield {
            'name':name,
            'link':link
        }
        yield response.follow(url=link,callback =self.parse_restaurant)


def parse_restaurant(self,response):
    name = response.xpath("//h1[@class='postcard__title postcard__title--claimed']/text()").get()
    website = response.xpath("(//a[@class='profile__website__link']/@href)[1]").get()
    address = response.xpath("(//address[@class='profile__address--compact']/text())[1]").get()

    yield{
        'name':name,
        "website":website,
        'address':address
    }
现有输出示例
2023-06-04 23:38:10 [scrapy.core.scraper] DEBUG: Scraped from <200 [https://zaubee.com/category/restaurant-in-fredonia-hclq6jom]>
{'name': 'Restaurants in Fredonia New York', 'link': '#'}

尝试跳转详情页时的输出:

2023-06-04 23:38:12 [scrapy.core.scraper] DEBUG: Scraped from <200 [https://zaubee.com/category/restaurant-in-fredonia-hclq6jom]>
{'name': None, 'website': None, 'address': None}
解决方案
  • 定位真实详情页链接:href为#说明跳转逻辑由JavaScript触发,真实链接通常存储在a标签的其他属性中(比如data-url、data-href)。打开浏览器开发者工具(F12),选中餐厅标题的a标签,查看其所有属性,找到存储真实详情页地址的字段。
  • 修改链接提取逻辑:将原代码中提取href的部分替换为提取真实链接的属性。例如,如果真实链接在data-url中,修改代码如下:
    link = restaurant.xpath(".//@data-url").get()
    
    若提取到的是相对路径,response.follow会自动拼接域名,无需额外处理;如果是完整URL,直接使用即可。
  • 修正详情页数据提取的XPath:当前输出None是因为跳转的页面并非真实详情页,或者XPath定位错误。拿到正确的详情页链接后,打开页面检查元素的实际class和结构,调整XPath表达式。例如:
    def parse_restaurant(self,response):
        # 清理文本中的空白字符
        name = response.xpath("//h1[@class='postcard__title']/text()").get(default='').strip()
        website = response.xpath("//a[@class='profile__website__link']/@href").get()
        address = response.xpath("//address[@class='profile__address--compact']/text()").get(default='').strip()
        # 新增电话、营业时间提取(根据实际页面结构调整)
        phone = response.xpath("//span[@class='profile__phone']/text()").get(default='').strip()
        hours = response.xpath("//div[@class='profile__hours']/text()").get(default='').strip()
    
        yield{
            'name':name,
            "website":website,
            'address':address,
            'phone':phone,
            'business_hours':hours
        }
    
  • 验证请求有效性:确保allowed_domains包含正确的域名(当前代码中是www.zaubee.com,但实际请求的是zaubee.com,可修改为allowed_domains = ['zaubee.com']避免被拦截)。

内容的提问来源于stack exchange,提问作者Licey Soremap

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 15:27:48