You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将栅格/矢量空间数据集转换为COCO格式用于目标检测?

实现地理坐标到像素坐标的转换,生成COCO格式数据集

完全可以实现这种坐标转换,核心是利用遥感影像的**地理变换参数(GeoTransform)**建立地理坐标与像素坐标的映射关系,结合Python的rasterio、geopandas等工具就能完成整个流程。以下是具体实现步骤和代码示例:

核心原理

遥感TIFF文件自带的地理变换参数(由rasterio读取的transform对象)定义了地理坐标(经纬度/投影坐标)到像素坐标的线性映射关系,rasterio提供的rowcol()方法可直接完成坐标转换计算,无需手动推导公式。

具体实现步骤

1. 读取遥感TIFF并导出为PNG

使用rasterio读取带地理参考的TIFF,获取影像尺寸和地理变换参数,再将影像数据导出为PNG格式:

import rasterio
from PIL import Image
import numpy as np

# 读取TIFF影像
with rasterio.open("satellite_image.tif") as src:
    img_width = src.width
    img_height = src.height
    # 获取地理变换参数(关键:用于坐标转换)
    geo_transform = src.transform
    # 读取影像数据(rasterio默认输出格式为 (波段数, 高度, 宽度))
    img_array = src.read()

# 将多波段影像转换为PNG兼容格式(以3波段RGB为例)
if img_array.shape[0] == 3:
    # 转成 (高度, 宽度, 3) 的RGB格式
    img_array = np.transpose(img_array, (1, 2, 0))
elif img_array.shape[0] == 4:
    # 如果是4波段(RGB+NIR),取前3波段转RGB
    img_array = np.transpose(img_array[:3, :, :], (1, 2, 0))

# 保存为PNG
img = Image.fromarray(img_array.astype(np.uint8))
img.save("satellite_image.png")

2. 读取地理标注并转换坐标系

使用geopandas读取SHP/GEOJSON格式的标注,确保标注的坐标系与TIFF影像一致(不一致则转换投影):

import geopandas as gpd

# 读取GeoJSON标注文件
gdf = gpd.read_file("annotations.geojson")

# 检查并匹配坐标系:确保标注与TIFF的CRS一致
with rasterio.open("satellite_image.tif") as src:
    tiff_crs = src.crs
if gdf.crs != tiff_crs:
    gdf = gdf.to_crs(tiff_crs)

3. 地理坐标转像素坐标,生成COCO格式标注

遍历每个标注的几何图形,用rasterio.transform.rowcol()将地理坐标转为像素坐标,再按照COCO格式组织数据:

import json
from shapely.geometry import Polygon, MultiPolygon

# 初始化COCO格式字典
coco_dataset = {
    "info": {},
    "licenses": [],
    "categories": [{"id": 1, "name": "target_object", "supercategory": "object"}],  # 根据你的类别修改
    "images": [
        {
            "id": 1,
            "width": img_width,
            "height": img_height,
            "file_name": "satellite_image.png",
            "license": 0,
            "date_captured": ""
        }
    ],
    "annotations": []
}

annotation_id = 1
for idx, row in gdf.iterrows():
    geom = row.geometry
    category_id = row.get("category_id", 1)  # 从标注文件中获取类别ID,无则默认1

    # 处理Polygon类型标注
    if isinstance(geom, Polygon):
        coords = list(geom.exterior.coords)
        pixel_coords = []
        for lon, lat in coords:
            # 转换地理坐标到像素坐标:返回(row, col),对应COCO的(y, x)
            y, x = rasterio.transform.rowcol(geo_transform, lon, lat)
            # 确保坐标在影像范围内,避免越界
            x = max(0, min(img_width - 1, x))
            y = max(0, min(img_height - 1, y))
            pixel_coords.extend([x, y])
        
        # 计算COCO格式的bbox(xmin, ymin, width, height)
        x_list = pixel_coords[::2]
        y_list = pixel_coords[1::2]
        xmin = min(x_list)
        ymin = min(y_list)
        bbox_w = max(x_list) - xmin
        bbox_h = max(y_list) - ymin

        # 添加标注到COCO数据集
        coco_dataset["annotations"].append({
            "id": annotation_id,
            "image_id": 1,
            "category_id": category_id,
            "segmentation": [pixel_coords],
            "area": round(geom.area, 2),  # 地理面积,或用像素面积:bbox_w * bbox_h
            "bbox": [xmin, ymin, bbox_w, bbox_h],
            "iscrowd": 0
        })
        annotation_id += 1
    
    # 处理MultiPolygon类型标注(拆分每个子Polygon)
    elif isinstance(geom, MultiPolygon):
        for sub_poly in geom.geoms:
            coords = list(sub_poly.exterior.coords)
            pixel_coords = []
            for lon, lat in coords:
                y, x = rasterio.transform.rowcol(geo_transform, lon, lat)
                x = max(0, min(img_width - 1, x))
                y = max(0, min(img_height - 1, y))
                pixel_coords.extend([x, y])
            
            x_list = pixel_coords[::2]
            y_list = pixel_coords[1::2]
            xmin = min(x_list)
            ymin = min(y_list)
            bbox_w = max(x_list) - xmin
            bbox_h = max(y_list) - ymin

            coco_dataset["annotations"].append({
                "id": annotation_id,
                "image_id": 1,
                "category_id": category_id,
                "segmentation": [pixel_coords],
                "area": round(sub_poly.area, 2),
                "bbox": [xmin, ymin, bbox_w, bbox_h],
                "iscrowd": 0
            })
            annotation_id += 1

# 保存COCO标注文件
with open("coco_annotations.json", "w", encoding="utf-8") as f:
    json.dump(coco_dataset, f, indent=2)

注意事项

  • 若影像有仿射变换(非正射),rasterio的rowcol()方法仍能正确处理,无需额外调整。
  • 像素坐标需确保在[0, width-1]和[0, height-1]范围内,避免标注超出图像边界。
  • COCO格式的segmentation要求是一维数组,每个多边形的坐标按[x1,y1,x2,y2,...]顺序排列。
  • 多类别标注需修改categories字段,确保category_id与标注文件中的类别对应。

内容的提问来源于stack exchange,提问作者Andres Camilo Zuñiga Gonzalez

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.21 11:22:30