请求排查爱沙尼亚卫星影像爬取代码问题并提供方案
爬取爱沙尼亚卫星影像代码排查求助
需要从链接https://xgis.maaamet.ee/xgis2/page/app/ristipuud爬取500张最新年份的tif格式卫星影像(该网站共约6000张影像,可通过分块编号从指定链接下载,以zip包形式提供)。现有Python爬取代码无报错但无法下载数据,作为编程新手,恳请帮忙排查代码错误并指导解决,附本人的爬取代码及立陶宛地区的参考爬取代码如下:
我的爬取代码
import re import requests from bs4 import BeautifulSoup webpage = 'https://xgis.maaamet.ee/xgis2/page/app/ristipuud' response = requests.get(site) # 错误:site变量未定义 bsoup = BeautifulSoup(response.text, 'html.parser') img_tags = soup.find_all('img') # 错误:soup未定义,应使用bsoup urls = [img['src'] for img in img_tags] for url in urls: filename = re.search(r'/([\w_-]+[.](jpg|gif|tif|png))$', url) if not filename: print("didn't match with the url: {}".format(url)) continue with open(filename.group(1), 'wb') as f: if 'http' not in url: url = '{}{}'.format(webpage, url) # 错误:链接拼接逻辑不符合网站规则 response = requests.get(url) f.write(response.content)
立陶宛地区参考代码
import time import requests from bs4 import BeautifulSoup import os def download_url(url, save_path, chunk_size=128): r = requests.get(url, stream=True) with open(save_path, 'wb') as fd: for chunk in r.iter_content(chunk_size=chunk_size): fd.write(chunk) def get_file_name(url): tokens = url.split("/") file_name = tokens[-1].split("?")[0] return file_name # Start timer start_time = time.time() print("Start time: ", start_time) # Create image directory image_directory = 'images' isExist = os.path.exists(image_directory) if not isExist: os.makedirs(image_directory) template = "https://www.geoportal.lt/map/webapp/rest/mapgateway/6100e156c755e15f6e46a8820824d8c595d30ae51?f=json" response = requests.get(template) if response.status_code == 200: soup = BeautifulSoup(response.content, "html.parser") link = soup.find("a") if link is not None: url = 'https://www.geoportal.lt/' + link['href'] file_name = get_file_name(url) print(file_name) # Save zip file download_url(url, './' + image_directory + '/' + file_name) # End timer end_time = time.time() # Calculate elapsed time elapsed_time = end_time - start_time print("Elapsed time: ", elapsed_time)
代码错误排查
变量问题
- 未定义
site变量,调用requests.get(site)会报错,应替换为已定义的webpage - BeautifulSoup对象命名为
bsoup,但后续查找标签用了soup.find_all('img'),变量名不匹配,需统一为bsoup
- 未定义
资源定位错误
- 目标网站的tif影像并非以
<img>标签形式存在,而是通过分块编号对应的API接口提供zip包下载,直接抓取<img>标签无法获取真实下载链接
- 目标网站的tif影像并非以
链接拼接错误
- 页面URL
webpage是前端页面地址,直接拼接相对路径无法得到正确的资源下载地址,需要先确认网站的下载API规则
- 页面URL
格式处理错误
- 目标资源是zip压缩包,代码直接将内容保存为tif文件,会导致文件损坏,需先下载zip包再解压提取tif文件
修正后的实现思路与示例代码
- 通过浏览器开发者工具查看网站的下载请求,获取真实的API地址和分块编号规则
- 筛选出最新年份的分块编号列表
- 批量下载zip包并解压提取tif文件
import requests import os import zipfile from tqdm import tqdm # 替换为网站真实的下载API地址 base_download_url = "https://xgis.maaamet.ee/xgis2/rest/ristipuud/download/" # 替换为从网站获取的最新年份500个分块编号 latest_block_ids = [f"tile_{i}" for i in range(1, 501)] # 创建保存目录 save_dir = "estonia_satellite_tifs" os.makedirs(save_dir, exist_ok=True) for block_id in tqdm(latest_block_ids, desc="批量下载中"): download_url = f"{base_download_url}{block_id}" try: # 流式请求适合大文件下载 response = requests.get(download_url, stream=True, timeout=30) response.raise_for_status() # 捕获HTTP请求错误 zip_file_path = os.path.join(save_dir, f"{block_id}.zip") # 写入zip包 with open(zip_file_path, "wb") as zip_f: for chunk in response.iter_content(chunk_size=8192): zip_f.write(chunk) # 解压zip包,仅提取tif文件 with zipfile.ZipFile(zip_file_path, 'r') as zip_ref: for file_name in zip_ref.namelist(): if file_name.lower().endswith(".tif"): zip_ref.extract(file_name, save_dir) # 可选:删除zip包节省空间 os.remove(zip_file_path) print(f"{block_id} 下载解压完成") except Exception as e: print(f"{block_id} 处理失败: {str(e)}")
注:需通过浏览器开发者工具抓包,获取网站真实的
base_download_url和最新年份的分块编号列表,替换示例中的对应内容。
内容的提问来源于stack exchange,提问作者Mizanur Rahman
相关产品推荐
相关产品推荐

