Scrapy爬虫改造:编写XPath/CSS选择器获取餐厅经纬度
Scrapy+Selenium爬虫获取餐厅经纬度解决方案
问题背景
我正在编写爬虫脚本,从网站https://pfchangsmexico.com.mx/ubicaciones/爬取餐厅数据,目前已实现获取餐厅名称、地址、邮编、电话等信息,但无法获取每个餐厅的经纬度。需要编写对应的XPath/CSS选择器,并整合到现有Scrapy+Selenium代码中。
原代码如下:
import scrapy import re from scrapy_selenium import SeleniumRequest from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium import webdriver class BistroSpider(scrapy.Spider): name = "bistro" allowed_domains = ["pfchangsmexico.com.mx"] start_urls = ["https://pfchangsmexico.com.mx/index.html"] """def __init__(self): super(BistroSpider, self).__init__() self.selenium = webdriver.Chrome()""" def parse(self, response): location_page = response.css('div a::attr(href)').get() yield SeleniumRequest(url=location_page, callback=self.parse_info) def parse_info(self, response): """iframe_locator = (By.XPATH, '//div/iframe') WebDriverWait(self.selenium, 10).until(EC.frame_to_be_available_and_switch_to_it(iframe_locator))""" res_names = response.xpath('//div/p[1]/span[1]/text()').getall() res_names = res_names[1:] print(res_names) print(len(res_names)) res_address = response.xpath('//div/p[1]/span[2]/text()').getall() print(res_address) print(len(res_address)) addresses = response.xpath('//div/p/span[2]').getall() addresses = [address for address in addresses if address != '<span style="font-family: Avenir;">Servicio a domicilio:</span>'] addresses = [address.replace('</span>', '') for address in addresses] addresses = [re.sub(r'<.*?>', '', address) for address in addresses] print(addresses) print(len(addresses)) postcode = response.xpath('//div/p[1]/span[3]/text()').getall() print(postcode) print(len(postcode)) phoneno = response.xpath('//div/p[4]/a[1]/span[1]/text()').getall() print(phoneno) print(len(phoneno)) """ Guadalajara (3) - 3rds path thats why 27 phone no in output //*[@id="guadalajara"]/div[2]/div[5]/div[1]/div/div/div[2]/div/p[3]/a[1]/span""" """iframe_element = self.selenium.find_element(By.XPATH, '//div/iframe') iframe_content = iframe_element.get_attribute("innerHTML") print("Iframe Content:") print(iframe_content)"""
关键分析
经纬度数据嵌入在页面的iframe(Google地图容器)中,必须先切换到iframe上下文才能访问内部元素。
定位与提取方法
- 切换到iframe:使用XPath定位地图iframe,等待其加载完成后切换上下文
iframe_locator = (By.XPATH, '//div[contains(@class, "map-container")]/iframe') WebDriverWait(response.request.meta['driver'], 10).until(EC.frame_to_be_available_and_switch_to_it(iframe_locator)) - 提取经纬度:每个餐厅标记的经纬度存储在
data-lat和data-lng属性中,对应的定位选择器:- 纬度XPath:
//div[@class="gm-style-moc"]//div[@data-lat]/@data-lat - 经度XPath:
//div[@class="gm-style-moc"]//div[@data-lng]/@data-lng
- 纬度XPath:
整合后的完整代码
import scrapy import re from scrapy_selenium import SeleniumRequest from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC class BistroSpider(scrapy.Spider): name = "bistro" allowed_domains = ["pfchangsmexico.com.mx"] start_urls = ["https://pfchangsmexico.com.mx/index.html"] def parse(self, response): # 精准定位到地址页面链接 location_page = response.css('div a[href*="ubicaciones"]::attr(href)').get() yield SeleniumRequest(url=location_page, callback=self.parse_info) def parse_info(self, response): # 获取scrapy-selenium管理的driver实例 driver = response.request.meta['driver'] lat_list = [] lng_list = [] # 切换到地图iframe并提取经纬度 try: iframe_locator = (By.XPATH, '//div[contains(@class, "map-container")]/iframe') WebDriverWait(driver, 10).until(EC.frame_to_be_available_and_switch_to_it(iframe_locator)) # 提取所有标记的经纬度属性 lat_elements = driver.find_elements(By.XPATH, '//div[@class="gm-style-moc"]//div[@data-lat]') lng_elements = driver.find_elements(By.XPATH, '//div[@class="gm-style-moc"]//div[@data-lng]') lat_list = [elem.get_attribute('data-lat') for elem in lat_elements] lng_list = [elem.get_attribute('data-lng') for elem in lng_elements] # 切回主页面上下文,避免后续元素定位出错 driver.switch_to.default_content() except Exception as e: self.logger.error(f"提取经纬度失败: {str(e)}") # 原有数据提取逻辑 res_names = response.xpath('//div/p[1]/span[1]/text()').getall()[1:] addresses = response.xpath('//div/p/span[2]').getall() addresses = [re.sub(r'<.*?>', '', addr) for addr in addresses if addr != '<span style="font-family: Avenir;">Servicio a domicilio:</span>'] postcode = response.xpath('//div/p[1]/span[3]/text()').getall() phoneno = response.xpath('//div/p[4]/a[1]/span[1]/text()').getall() # 整合所有数据并输出 for idx, name in enumerate(res_names): yield { 'name': name.strip(), 'address': addresses[idx].strip() if idx < len(addresses) else '', 'postcode': postcode[idx].strip() if idx < len(postcode) else '', 'phone': phoneno[idx].strip() if idx < len(phoneno) else '', 'latitude': lat_list[idx] if idx < len(lat_list) else '', 'longitude': lng_list[idx] if idx < len(lng_list) else '' }
注意事项
- 无需手动初始化driver,scrapy-selenium会自动管理driver实例,通过
response.request.meta['driver']获取即可 - 提取完iframe内数据后,必须切回主页面上下文,否则后续主页面元素定位会失败
- 如果经纬度数量与餐厅数量不匹配,可调整iframe内标记的XPath规则,确保定位到每个餐厅对应的地图标记
内容的提问来源于stack exchange,提问作者Prathamesh
相关产品推荐
相关产品推荐

