You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Scrapy爬虫改造:编写XPath/CSS选择器获取餐厅经纬度

Scrapy+Selenium爬虫获取餐厅经纬度解决方案

问题背景

我正在编写爬虫脚本,从网站https://pfchangsmexico.com.mx/ubicaciones/爬取餐厅数据,目前已实现获取餐厅名称、地址、邮编、电话等信息,但无法获取每个餐厅的经纬度。需要编写对应的XPath/CSS选择器,并整合到现有Scrapy+Selenium代码中。

原代码如下:

import scrapy
import re
from scrapy_selenium import SeleniumRequest
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium import webdriver

class BistroSpider(scrapy.Spider):
    name = "bistro"
    allowed_domains = ["pfchangsmexico.com.mx"]
    start_urls = ["https://pfchangsmexico.com.mx/index.html"]

    """def __init__(self):
        super(BistroSpider, self).__init__()
        self.selenium = webdriver.Chrome()"""
    
    def parse(self, response):
        location_page = response.css('div a::attr(href)').get()
        yield SeleniumRequest(url=location_page, callback=self.parse_info)
    
    def parse_info(self, response):

        """iframe_locator = (By.XPATH, '//div/iframe')
        WebDriverWait(self.selenium, 10).until(EC.frame_to_be_available_and_switch_to_it(iframe_locator))"""

        
        res_names = response.xpath('//div/p[1]/span[1]/text()').getall()
        res_names = res_names[1:]
        print(res_names)
        print(len(res_names))
        res_address = response.xpath('//div/p[1]/span[2]/text()').getall()
        print(res_address)
        print(len(res_address))

        addresses = response.xpath('//div/p/span[2]').getall()

        addresses = [address for address in addresses if address !=
                     '<span style="font-family: Avenir;">Servicio a domicilio:</span>']
        addresses = [address.replace('</span>', '') for address in addresses]
        addresses = [re.sub(r'<.*?>', '', address) for address in addresses]

        print(addresses)
        print(len(addresses))

        postcode = response.xpath('//div/p[1]/span[3]/text()').getall()
        print(postcode)
        print(len(postcode))

        phoneno = response.xpath('//div/p[4]/a[1]/span[1]/text()').getall()
        print(phoneno)
        print(len(phoneno))

        """ Guadalajara (3) - 3rds path thats why 27 phone no in output
        //*[@id="guadalajara"]/div[2]/div[5]/div[1]/div/div/div[2]/div/p[3]/a[1]/span"""

        """iframe_element = self.selenium.find_element(By.XPATH, '//div/iframe')
        iframe_content = iframe_element.get_attribute("innerHTML")
        print("Iframe Content:")
        print(iframe_content)"""

关键分析

经纬度数据嵌入在页面的iframe(Google地图容器)中,必须先切换到iframe上下文才能访问内部元素。

定位与提取方法

  1. 切换到iframe:使用XPath定位地图iframe,等待其加载完成后切换上下文
    iframe_locator = (By.XPATH, '//div[contains(@class, "map-container")]/iframe')
    WebDriverWait(response.request.meta['driver'], 10).until(EC.frame_to_be_available_and_switch_to_it(iframe_locator))
    
  2. 提取经纬度:每个餐厅标记的经纬度存储在data-lat和data-lng属性中,对应的定位选择器:
    • 纬度XPath://div[@class="gm-style-moc"]//div[@data-lat]/@data-lat
    • 经度XPath://div[@class="gm-style-moc"]//div[@data-lng]/@data-lng

整合后的完整代码

import scrapy
import re
from scrapy_selenium import SeleniumRequest
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

class BistroSpider(scrapy.Spider):
    name = "bistro"
    allowed_domains = ["pfchangsmexico.com.mx"]
    start_urls = ["https://pfchangsmexico.com.mx/index.html"]
    
    def parse(self, response):
        # 精准定位到地址页面链接
        location_page = response.css('div a[href*="ubicaciones"]::attr(href)').get()
        yield SeleniumRequest(url=location_page, callback=self.parse_info)
    
    def parse_info(self, response):
        # 获取scrapy-selenium管理的driver实例
        driver = response.request.meta['driver']
        lat_list = []
        lng_list = []
        
        # 切换到地图iframe并提取经纬度
        try:
            iframe_locator = (By.XPATH, '//div[contains(@class, "map-container")]/iframe')
            WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it(iframe_locator))
            
            # 提取所有标记的经纬度属性
            lat_elements = driver.find_elements(By.XPATH, '//div[@class="gm-style-moc"]//div[@data-lat]')
            lng_elements = driver.find_elements(By.XPATH, '//div[@class="gm-style-moc"]//div[@data-lng]')
            
            lat_list = [elem.get_attribute('data-lat') for elem in lat_elements]
            lng_list = [elem.get_attribute('data-lng') for elem in lng_elements]
            
            # 切回主页面上下文,避免后续元素定位出错
            driver.switch_to.default_content()
        except Exception as e:
            self.logger.error(f"提取经纬度失败: {str(e)}")
        
        # 原有数据提取逻辑
        res_names = response.xpath('//div/p[1]/span[1]/text()').getall()[1:]
        addresses = response.xpath('//div/p/span[2]').getall()
        addresses = [re.sub(r'<.*?>', '', addr) for addr in addresses if addr != '<span style="font-family: Avenir;">Servicio a domicilio:</span>']
        postcode = response.xpath('//div/p[1]/span[3]/text()').getall()
        phoneno = response.xpath('//div/p[4]/a[1]/span[1]/text()').getall()
        
        # 整合所有数据并输出
        for idx, name in enumerate(res_names):
            yield {
                'name': name.strip(),
                'address': addresses[idx].strip() if idx < len(addresses) else '',
                'postcode': postcode[idx].strip() if idx < len(postcode) else '',
                'phone': phoneno[idx].strip() if idx < len(phoneno) else '',
                'latitude': lat_list[idx] if idx < len(lat_list) else '',
                'longitude': lng_list[idx] if idx < len(lng_list) else ''
            }

注意事项

  • 无需手动初始化driver,scrapy-selenium会自动管理driver实例,通过response.request.meta['driver']获取即可
  • 提取完iframe内数据后,必须切回主页面上下文,否则后续主页面元素定位会失败
  • 如果经纬度数量与餐厅数量不匹配,可调整iframe内标记的XPath规则,确保定位到每个餐厅对应的地图标记

内容的提问来源于stack exchange,提问作者Prathamesh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.16 12:15:19