You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Web Scraping新手求助:如何从目标网页提取经纬度

从urbania.pe网页提取经纬度的修正方案

你的代码存在两个核心问题导致提取失败,以下是问题分析和修正后的实现:

问题分析

  • BeautifulSoup用法错误:tag.find('latitude')是在HTML标签中查找<latitude>元素,但latitude是JavaScript对象的属性,并非HTML标签,因此永远返回None。
  • 正则参数类型错误:re.findall需要传入字符串作为匹配源,但你传入的是BeautifulSoup的ResultSet对象(标签集合),无法直接匹配。

修正代码

import requests
import re
import json
import urllib.parse

sa_key = 'ea69223fa47f72fac0907759'
sa_api = 'https://api.scrapingant.com/v2/general'

page = 'https://urbania.pe/inmueble/proyecto/ememhvin-proyecto-mariscal-castilla-lima-santiago-de-surco-tale-inmobiliaria-65659522'

qParams = {'url': page, 'x-api-key': sa_key}
reqUrl = f'{sa_api}?{urllib.parse.urlencode(qParams)}'

r = requests.get(reqUrl)
# 筛选包含POSTING和postingGeolocation的脚本内容
target_script = None
for script in re.findall(r'<script type="text/javascript">(.*?)</script>', r.text, re.DOTALL):
    if 'POSTING' in script and 'postingGeolocation' in script:
        target_script = script
        break

if target_script:
    # 匹配geolocation的JSON结构
    geo_match = re.search(r'postingGeolocation\.geolocation\s*=\s*({.*?});', target_script, re.DOTALL)
    if geo_match:
        geo_dict = json.loads(geo_match.group(1))
        lat = geo_dict.get('latitude')
        lon = geo_dict.get('longitude')
        print(f"纬度:{lat},经度:{lon}")
    else:
        print("未找到geolocation数据块")
else:
    print("未定位到包含目标数据的脚本标签")

代码说明

  1. 直接从响应文本中筛选包含POSTING对象的脚本,避免BeautifulSoup遍历所有标签的冗余操作
  2. 用正则精准匹配postingGeolocation.geolocation = {...}的结构,提取其中的JSON字符串
  3. 通过json.loads解析JSON,直接获取经纬度属性,比纯正则匹配数值更稳定可靠

内容的提问来源于stack exchange,提问作者DANIEL HERNANDEZ

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 17:21:06