You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Scrapy爬取yaencontre巴塞罗那租房信息:经纬度获取失败求助

解决Yaencontre公寓经纬度爬取失败问题

问题分析

你当前代码提取经纬度失败的核心原因:

  1. XPath使用extract()返回列表而非单个值,无法直接获取有效数据;
  2. XPath表达式中用*1做类型转换无效,XPath 1.0不支持算术运算;
  3. 依赖图片src字符串截取的方式稳定性差,页面链接参数结构变化就会失效。

修复方案

以下提供两种可靠提取方式,优先推荐方案二(结构化数据提取,稳定性更强):

方案一:修复图片src提取逻辑

用正则匹配图片链接中的经纬度参数,替代易出错的XPath字符串截取:

#!/usr/local/bin/env python3
import scrapy
from scraping.items import fields
import re

n = 'barcelona'
list_of_urls = []
for i in range(1,2):
    url = 'https://www.yaencontre.com/alquiler/pisos/barcelona/pag-{}'.format(i)
    list_of_urls.append(url)

class scraperApp(scrapy.Spider):
    name = n
    start_urls = list_of_urls
    
    def parse(self,response):
        for href in response.xpath("//a[@class= 'd-ellipsis']/@href"):
            u = 'https://www.yaencontre.com'+ href.extract()    
            yield scrapy.Request(u, callback=self.parse_dir_contents)                 

    def parse_dir_contents(self,response):
        item = fields() 
        item['vivienda'] = n
        
        # 提取并清理价格数据
        price_text = response.xpath("//div[@class='price-wrapper mb-sm']/span/text()").extract_first()
        item['price'] = price_text.strip() if price_text else None
        
        # 从地图图片链接提取经纬度
        map_src = response.xpath("//img[@class='d-block']/@src").extract_first()
        if map_src:
            lat_lon_match = re.search(r'center=([\d.]+)%2C([\d.]+)', map_src)
            if lat_lon_match:
                item['lat'] = float(lat_lon_match.group(1))
                item['lon'] = float(lat_lon_match.group(2))
            else:
                item['lat'] = item['lon'] = None
        else:
            item['lat'] = item['lon'] = None
            
        yield item  

方案二:提取页面结构化数据(推荐)

网站会嵌入JSON-LD格式的结构化数据,包含标准的经纬度信息,这种方式不受页面布局变化影响:

#!/usr/local/bin/env python3
import scrapy
from scraping.items import fields
import json

n = 'barcelona'
list_of_urls = []
for i in range(1,2):
    url = 'https://www.yaencontre.com/alquiler/pisos/barcelona/pag-{}'.format(i)
    list_of_urls.append(url)

class scraperApp(scrapy.Spider):
    name = n
    start_urls = list_of_urls
    
    def parse(self,response):
        for href in response.xpath("//a[@class= 'd-ellipsis']/@href"):
            u = 'https://www.yaencontre.com'+ href.extract()    
            yield scrapy.Request(u, callback=self.parse_dir_contents)                 

    def parse_dir_contents(self,response):
        item = fields() 
        item['vivienda'] = n
        
        # 提取并清理价格数据
        price_text = response.xpath("//div[@class='price-wrapper mb-sm']/span/text()").extract_first()
        item['price'] = price_text.strip() if price_text else None
        
        # 从JSON-LD结构化数据提取经纬度
        json_ld_text = response.xpath("//script[@type='application/ld+json']/text()").extract_first()
        if json_ld_text:
            try:
                json_data = json.loads(json_ld_text)
                # 处理JSON-LD可能是列表的情况
                if isinstance(json_data, list):
                    json_data = json_data[0]
                geo_info = json_data.get('geo')
                if geo_info:
                    item['lat'] = float(geo_info.get('latitude'))
                    item['lon'] = float(geo_info.get('longitude'))
                else:
                    item['lat'] = item['lon'] = None
            except json.JSONDecodeError:
                item['lat'] = item['lon'] = None
        else:
            item['lat'] = item['lon'] = None
            
        yield item  

额外优化点

  • 简化了list_of_urls生成逻辑,去除冗余循环;
  • 增加空值容错处理,避免爬虫因页面结构变化崩溃;
  • 价格提取仅保留文本内容并清理,避免返回HTML标签。

内容的提问来源于stack exchange,提问作者warforterritory

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.27 04:37:52