You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

请求排查爱沙尼亚卫星影像爬取代码问题并提供方案

爬取爱沙尼亚卫星影像代码排查求助

需要从链接https://xgis.maaamet.ee/xgis2/page/app/ristipuud爬取500张最新年份的tif格式卫星影像(该网站共约6000张影像,可通过分块编号从指定链接下载,以zip包形式提供)。现有Python爬取代码无报错但无法下载数据,作为编程新手,恳请帮忙排查代码错误并指导解决,附本人的爬取代码及立陶宛地区的参考爬取代码如下:

我的爬取代码

import re
import requests
from bs4 import BeautifulSoup

webpage = 'https://xgis.maaamet.ee/xgis2/page/app/ristipuud'

response = requests.get(site)  # 错误:site变量未定义
bsoup = BeautifulSoup(response.text, 'html.parser')
img_tags = soup.find_all('img')  # 错误:soup未定义,应使用bsoup

urls = [img['src'] for img in img_tags]

for url in urls:
    filename = re.search(r'/([\w_-]+[.](jpg|gif|tif|png))$', url)
    if not filename:
        print("didn't match with the url: {}".format(url))
        continue
    with open(filename.group(1), 'wb') as f:
        if 'http' not in url:
            url = '{}{}'.format(webpage, url)  # 错误:链接拼接逻辑不符合网站规则
        response = requests.get(url)
        f.write(response.content)

立陶宛地区参考代码

import time
import requests
from bs4 import BeautifulSoup
import os

def download_url(url, save_path, chunk_size=128):
    r = requests.get(url, stream=True)
    with open(save_path, 'wb') as fd:
        for chunk in r.iter_content(chunk_size=chunk_size):
            fd.write(chunk)

def get_file_name(url):
    tokens = url.split("/")
    file_name = tokens[-1].split("?")[0]
    return file_name

# Start timer
start_time = time.time()
print("Start time: ", start_time)

# Create image directory
image_directory = 'images'
isExist = os.path.exists(image_directory)
if not isExist:
    os.makedirs(image_directory)

template = "https://www.geoportal.lt/map/webapp/rest/mapgateway/6100e156c755e15f6e46a8820824d8c595d30ae51?f=json"

response = requests.get(template)
if response.status_code == 200:
    soup = BeautifulSoup(response.content, "html.parser")
    link = soup.find("a")
    if link is not None:
        url = 'https://www.geoportal.lt/' + link['href']
        file_name = get_file_name(url)
        print(file_name)
        # Save zip file
        download_url(url, './' + image_directory + '/' + file_name)

# End timer
end_time = time.time()

# Calculate elapsed time
elapsed_time = end_time - start_time
print("Elapsed time: ", elapsed_time)

代码错误排查

  1. 变量问题

    • 未定义site变量,调用requests.get(site)会报错,应替换为已定义的webpage
    • BeautifulSoup对象命名为bsoup,但后续查找标签用了soup.find_all('img'),变量名不匹配,需统一为bsoup
  2. 资源定位错误

    • 目标网站的tif影像并非以<img>标签形式存在,而是通过分块编号对应的API接口提供zip包下载,直接抓取<img>标签无法获取真实下载链接
  3. 链接拼接错误

    • 页面URLwebpage是前端页面地址,直接拼接相对路径无法得到正确的资源下载地址,需要先确认网站的下载API规则
  4. 格式处理错误

    • 目标资源是zip压缩包,代码直接将内容保存为tif文件,会导致文件损坏,需先下载zip包再解压提取tif文件

修正后的实现思路与示例代码

  1. 通过浏览器开发者工具查看网站的下载请求,获取真实的API地址和分块编号规则
  2. 筛选出最新年份的分块编号列表
  3. 批量下载zip包并解压提取tif文件
import requests
import os
import zipfile
from tqdm import tqdm

# 替换为网站真实的下载API地址
base_download_url = "https://xgis.maaamet.ee/xgis2/rest/ristipuud/download/"
# 替换为从网站获取的最新年份500个分块编号
latest_block_ids = [f"tile_{i}" for i in range(1, 501)]

# 创建保存目录
save_dir = "estonia_satellite_tifs"
os.makedirs(save_dir, exist_ok=True)

for block_id in tqdm(latest_block_ids, desc="批量下载中"):
    download_url = f"{base_download_url}{block_id}"
    try:
        # 流式请求适合大文件下载
        response = requests.get(download_url, stream=True, timeout=30)
        response.raise_for_status()  # 捕获HTTP请求错误
        
        zip_file_path = os.path.join(save_dir, f"{block_id}.zip")
        # 写入zip包
        with open(zip_file_path, "wb") as zip_f:
            for chunk in response.iter_content(chunk_size=8192):
                zip_f.write(chunk)
        
        # 解压zip包,仅提取tif文件
        with zipfile.ZipFile(zip_file_path, 'r') as zip_ref:
            for file_name in zip_ref.namelist():
                if file_name.lower().endswith(".tif"):
                    zip_ref.extract(file_name, save_dir)
        
        # 可选:删除zip包节省空间
        os.remove(zip_file_path)
        print(f"{block_id} 下载解压完成")
    
    except Exception as e:
        print(f"{block_id} 处理失败: {str(e)}")

注:需通过浏览器开发者工具抓包,获取网站真实的base_download_url和最新年份的分块编号列表,替换示例中的对应内容。

内容的提问来源于stack exchange,提问作者Mizanur Rahman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 05:09:53