解决Flickr特定标签地理标签导出Shapefile时重复点过多的问题求助
解决Flickr特定标签地理标签导出Shapefile时重复点过多的问题求助
Hey 👋,我看了你的问题和代码,40万条记录去重后只剩4千,确实说明大量重复的照片被拉取了。先帮你分析下可能的问题,再给你修改方案:
可能的重复原因
- Flickr API的返回特性:当你用多语言同义词标签(比如同时搜"Alps"和"Alpen")时,同一照片如果被标记了多个你指定的标签,可能会在不同分页甚至同一请求里被重复返回。
- 代码缺少去重逻辑:你的代码只是把所有请求到的照片直接追加到列表,完全没做去重处理,哪怕是同一
photo["id"]的内容也会被多次添加。
修改后的代码(增加去重逻辑)
我主要做了两处核心修改,同时优化了异常处理:
- 用字典存储照片,以
photo["id"]作为唯一键,自动覆盖重复的同ID内容 - 增强了请求异常处理,避免单页出错导致整个流程中断
import requests import time import geopandas as gpd from shapely.geometry import Point API_KEY = "de874d66a7e115fe8d805b9630a246ce" TAGS = ["Alpen", "Alpi", "Alpes", "Alps", "Alpe", "Альпы", "阿尔卑斯山", "アルプス山脈", "جبال الألب", "आल्प्स ", "آلپس"] PER_PAGE = 500 def fetch_flickr_geotags(api_key, tags, per_page): url = "https://api.flickr.com/services/rest/" # 用字典存储,key为photo id,自动实现去重 all_photos = {} page = 1 tags_string = ",".join(tags) params = { "method": "flickr.photos.search", "api_key": api_key, "tags": tags_string, "tag_mode": "any", "has_geo": 1, "format": "json", "nojsoncallback": 1, "per_page": per_page, "page": page, "extras": "geo" } # 先获取总页数,避免无效循环 response = requests.get(url, params=params) data = response.json() if data["stat"] != "ok": raise Exception(f"Flickr API Fehler: {data['message']}") total_pages = data["photos"]["pages"] print(f"Gesamtzahl der Seiten: {total_pages}") while page <= total_pages: params["page"] = page try: response = requests.get(url, params=params) response.raise_for_status() # 检查HTTP请求是否成功 data = response.json() except requests.exceptions.RequestException as e: print(f"Seite {page} konnte nicht geladen werden: {str(e)}") time.sleep(2) # 出错后多等2秒再重试 page += 1 continue if "stat" in data and data["stat"] != "ok": print(f"API-Fehler auf Seite {page}: {data['message']}") page += 1 continue photos = data["photos"]["photo"] # 字典去重:同ID的照片只会保留最后一次获取的(内容基本一致) for photo in photos: all_photos[photo["id"]] = photo print(f"Seite {page} von {total_pages} heruntergeladen. Aktuell gespeicherte eindeutige Fotos: {len(all_photos)}") time.sleep(1) page += 1 # 把字典的值转成列表返回 return list(all_photos.values()) def save_photos_as_shapefile(photos, filename): records = [] for photo in photos: if "latitude" in photo and "longitude" in photo: try: lat = float(photo["latitude"]) lon = float(photo["longitude"]) point = Point(lon, lat) record = { "id": photo["id"], "title": photo["title"], "geometry": point } records.append(record) except ValueError: print(f"Ungültige Koordinaten für Foto {photo['id']}: {photo.get('latitude')}, {photo.get('longitude')}") continue if not records: print("Keine Geotags gefunden.") return gdf = gpd.GeoDataFrame(records, crs="EPSG:4326") gdf.to_file(filename) print(f"Shapefile gespeichert: {filename}") print(f"Gesamtzahl eindeutiger Punkte: {len(gdf)}") photos = fetch_flickr_geotags(API_KEY, TAGS, PER_PAGE) if photos: save_photos_as_shapefile(photos, "flickr_tags_shapefile.shp") else: print("Keine Fotos mit Geotags gefunden.")
额外优化建议
- 限定搜索范围:如果只需要阿尔卑斯山区的照片,可以加
bbox参数指定经纬度范围,进一步减少无关和重复结果 - 调整请求间隔:如果遇到API限流提示,可以把
time.sleep(1)改成2秒,避免触发Flickr的请求频率限制 - 验证坐标合理性:可以再加一层判断,确保坐标在阿尔卑斯山的大致范围内(比如纬度43-48,经度5-15),过滤掉错误标记的坐标
备注:内容来源于stack exchange,提问作者PcPrincipal
相关产品推荐
相关产品推荐

