You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

解决Flickr特定标签地理标签导出Shapefile时重复点过多的问题求助

解决Flickr特定标签地理标签导出Shapefile时重复点过多的问题求助

Hey 👋,我看了你的问题和代码,40万条记录去重后只剩4千,确实说明大量重复的照片被拉取了。先帮你分析下可能的问题,再给你修改方案:

可能的重复原因

  • Flickr API的返回特性:当你用多语言同义词标签(比如同时搜"Alps"和"Alpen")时,同一照片如果被标记了多个你指定的标签,可能会在不同分页甚至同一请求里被重复返回。
  • 代码缺少去重逻辑:你的代码只是把所有请求到的照片直接追加到列表,完全没做去重处理,哪怕是同一photo["id"]的内容也会被多次添加。

修改后的代码(增加去重逻辑)

我主要做了两处核心修改,同时优化了异常处理:

  1. 用字典存储照片,以photo["id"]作为唯一键,自动覆盖重复的同ID内容
  2. 增强了请求异常处理,避免单页出错导致整个流程中断
import requests
import time
import geopandas as gpd
from shapely.geometry import Point

API_KEY = "de874d66a7e115fe8d805b9630a246ce"
TAGS = ["Alpen", "Alpi", "Alpes", "Alps", "Alpe", "Альпы", "阿尔卑斯山", "アルプス山脈", "جبال الألب", "आल्प्स ", "آلپس"]
PER_PAGE = 500 

def fetch_flickr_geotags(api_key, tags, per_page):
    url = "https://api.flickr.com/services/rest/"
    # 用字典存储,key为photo id,自动实现去重
    all_photos = {}
    page = 1

    tags_string = ",".join(tags)

    params = {
        "method": "flickr.photos.search",
        "api_key": api_key,
        "tags": tags_string,  
        "tag_mode": "any",  
        "has_geo": 1,  
        "format": "json",
        "nojsoncallback": 1,
        "per_page": per_page,
        "page": page,
        "extras": "geo"
    }

    # 先获取总页数,避免无效循环
    response = requests.get(url, params=params)
    data = response.json()

    if data["stat"] != "ok":
        raise Exception(f"Flickr API Fehler: {data['message']}")

    total_pages = data["photos"]["pages"]
    print(f"Gesamtzahl der Seiten: {total_pages}")

    while page <= total_pages:
        params["page"] = page
        try:
            response = requests.get(url, params=params)
            response.raise_for_status()  # 检查HTTP请求是否成功
            data = response.json()
        except requests.exceptions.RequestException as e:
            print(f"Seite {page} konnte nicht geladen werden: {str(e)}")
            time.sleep(2)  # 出错后多等2秒再重试
            page += 1
            continue

        if "stat" in data and data["stat"] != "ok":
            print(f"API-Fehler auf Seite {page}: {data['message']}")
            page += 1
            continue

        photos = data["photos"]["photo"]
        # 字典去重:同ID的照片只会保留最后一次获取的(内容基本一致)
        for photo in photos:
            all_photos[photo["id"]] = photo

        print(f"Seite {page} von {total_pages} heruntergeladen. Aktuell gespeicherte eindeutige Fotos: {len(all_photos)}")
        time.sleep(1)  
        page += 1

    # 把字典的值转成列表返回
    return list(all_photos.values())

def save_photos_as_shapefile(photos, filename):
    records = []
    
    for photo in photos:
        if "latitude" in photo and "longitude" in photo:
            try:
                lat = float(photo["latitude"])
                lon = float(photo["longitude"])
                point = Point(lon, lat)
                record = {
                    "id": photo["id"],
                    "title": photo["title"],
                    "geometry": point
                }
                records.append(record)
            except ValueError:
                print(f"Ungültige Koordinaten für Foto {photo['id']}: {photo.get('latitude')}, {photo.get('longitude')}")
                continue

    if not records:
        print("Keine Geotags gefunden.")
        return

    gdf = gpd.GeoDataFrame(records, crs="EPSG:4326")
    gdf.to_file(filename)
    print(f"Shapefile gespeichert: {filename}")
    print(f"Gesamtzahl eindeutiger Punkte: {len(gdf)}")

photos = fetch_flickr_geotags(API_KEY, TAGS, PER_PAGE)

if photos:
    save_photos_as_shapefile(photos, "flickr_tags_shapefile.shp")
else:
    print("Keine Fotos mit Geotags gefunden.")

额外优化建议

  • 限定搜索范围:如果只需要阿尔卑斯山区的照片,可以加bbox参数指定经纬度范围,进一步减少无关和重复结果
  • 调整请求间隔:如果遇到API限流提示,可以把time.sleep(1)改成2秒,避免触发Flickr的请求频率限制
  • 验证坐标合理性:可以再加一层判断,确保坐标在阿尔卑斯山的大致范围内(比如纬度43-48,经度5-15),过滤掉错误标记的坐标

备注:内容来源于stack exchange,提问作者PcPrincipal

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.13 18:54:29