使用Scrapy爬取yaencontre巴塞罗那租房信息:经纬度获取失败求助
解决Yaencontre公寓经纬度爬取失败问题
问题分析
你当前代码提取经纬度失败的核心原因:
- XPath使用
extract()返回列表而非单个值,无法直接获取有效数据; - XPath表达式中用
*1做类型转换无效,XPath 1.0不支持算术运算; - 依赖图片src字符串截取的方式稳定性差,页面链接参数结构变化就会失效。
修复方案
以下提供两种可靠提取方式,优先推荐方案二(结构化数据提取,稳定性更强):
方案一:修复图片src提取逻辑
用正则匹配图片链接中的经纬度参数,替代易出错的XPath字符串截取:
#!/usr/local/bin/env python3 import scrapy from scraping.items import fields import re n = 'barcelona' list_of_urls = [] for i in range(1,2): url = 'https://www.yaencontre.com/alquiler/pisos/barcelona/pag-{}'.format(i) list_of_urls.append(url) class scraperApp(scrapy.Spider): name = n start_urls = list_of_urls def parse(self,response): for href in response.xpath("//a[@class= 'd-ellipsis']/@href"): u = 'https://www.yaencontre.com'+ href.extract() yield scrapy.Request(u, callback=self.parse_dir_contents) def parse_dir_contents(self,response): item = fields() item['vivienda'] = n # 提取并清理价格数据 price_text = response.xpath("//div[@class='price-wrapper mb-sm']/span/text()").extract_first() item['price'] = price_text.strip() if price_text else None # 从地图图片链接提取经纬度 map_src = response.xpath("//img[@class='d-block']/@src").extract_first() if map_src: lat_lon_match = re.search(r'center=([\d.]+)%2C([\d.]+)', map_src) if lat_lon_match: item['lat'] = float(lat_lon_match.group(1)) item['lon'] = float(lat_lon_match.group(2)) else: item['lat'] = item['lon'] = None else: item['lat'] = item['lon'] = None yield item
方案二:提取页面结构化数据(推荐)
网站会嵌入JSON-LD格式的结构化数据,包含标准的经纬度信息,这种方式不受页面布局变化影响:
#!/usr/local/bin/env python3 import scrapy from scraping.items import fields import json n = 'barcelona' list_of_urls = [] for i in range(1,2): url = 'https://www.yaencontre.com/alquiler/pisos/barcelona/pag-{}'.format(i) list_of_urls.append(url) class scraperApp(scrapy.Spider): name = n start_urls = list_of_urls def parse(self,response): for href in response.xpath("//a[@class= 'd-ellipsis']/@href"): u = 'https://www.yaencontre.com'+ href.extract() yield scrapy.Request(u, callback=self.parse_dir_contents) def parse_dir_contents(self,response): item = fields() item['vivienda'] = n # 提取并清理价格数据 price_text = response.xpath("//div[@class='price-wrapper mb-sm']/span/text()").extract_first() item['price'] = price_text.strip() if price_text else None # 从JSON-LD结构化数据提取经纬度 json_ld_text = response.xpath("//script[@type='application/ld+json']/text()").extract_first() if json_ld_text: try: json_data = json.loads(json_ld_text) # 处理JSON-LD可能是列表的情况 if isinstance(json_data, list): json_data = json_data[0] geo_info = json_data.get('geo') if geo_info: item['lat'] = float(geo_info.get('latitude')) item['lon'] = float(geo_info.get('longitude')) else: item['lat'] = item['lon'] = None except json.JSONDecodeError: item['lat'] = item['lon'] = None else: item['lat'] = item['lon'] = None yield item
额外优化点
- 简化了
list_of_urls生成逻辑,去除冗余循环; - 增加空值容错处理,避免爬虫因页面结构变化崩溃;
- 价格提取仅保留文本内容并清理,避免返回HTML标签。
内容的提问来源于stack exchange,提问作者warforterritory
相关产品推荐
相关产品推荐

