You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何遍历CSV生成的DataFrame,执行点在多边形操作并新增结果列

解决DataFrame遍历与点在多边形结果赋值的方案

不用手动逐行遍历DataFrame啦,用pandas的apply方法就能高效完成这个任务,结合你已有的shapely和fiona代码,我给你整理了完整的实现步骤:

步骤1:准备工作(读取数据+预加载多边形)

首先读取你的CSV数据到DataFrame,同时预加载GeoJSON里的所有多边形要素(提前转换为shapely对象,避免重复IO操作):

import pandas as pd
from shapely.geometry import Point, shape
import fiona

# 替换为你的CSV文件路径
df = pd.read_csv("your_input_data.csv")

# 预加载GeoJSON中的多边形及属性
fc = fiona.open("ngaadmbndaadm2osgof20170222.geojson")
polygon_list = []
for feature in fc:
    # 将GeoJSON要素转为shapely多边形
    poly_shape = shape(feature['geometry'])
    # 这里可以按需存储多边形的属性,比如行政区名称(根据你的GeoJSON schema调整键名)
    polygon_list.append({
        "shape": poly_shape,
        "admin_name": feature['properties'].get("ADM2_NAME")  # 替换为你实际需要的属性字段
    })
fc.close()  # 记得关闭文件句柄

步骤2:编写点在多边形的判断函数

根据你的需求,写一个函数接收经纬度,返回对应的多边形信息(或布尔值):

def find_polygon_for_point(lon, lat):
    point = Point(lon, lat)
    # 遍历所有多边形,判断点是否在内部
    for poly in polygon_list:
        if point.within(poly["shape"]):
            return poly["admin_name"]  # 返回多边形的名称,也可以返回True/False
    return None  # 如果点不在任何多边形内,返回None

步骤3:批量处理DataFrame,添加新列

用apply方法对DataFrame的每一行执行判断函数,直接生成新列:

# 假设你的CSV里经度列是"longitude",纬度列是"latitude",替换为实际列名
df["admin_region"] = df.apply(
    lambda row: find_polygon_for_point(row["longitude"], row["latitude"]),
    axis=1  # 指定按行处理
)

优化建议(针对大量多边形场景)

如果你的GeoJSON里有上百个甚至更多多边形,上面的遍历会比较慢,可以用空间索引(比如rtree库)来加速查询:

from rtree import index

# 创建空间索引,存储每个多边形的边界
spatial_idx = index.Index()
for idx, poly in enumerate(polygon_list):
    spatial_idx.insert(idx, poly["shape"].bounds)

def find_polygon_fast(lon, lat):
    point = Point(lon, lat)
    # 先通过空间索引筛选出可能包含点的多边形,减少遍历数量
    for idx in spatial_idx.intersection(point.bounds):
        if point.within(polygon_list[idx]["shape"]):
            return polygon_list[idx]["admin_name"]
    return None

# 用优化后的函数生成新列
df["admin_region_fast"] = df.apply(
    lambda row: find_polygon_fast(row["longitude"], row["latitude"]),
    axis=1
)

为什么推荐用apply而不是手动遍历?因为pandas的apply是内部优化过的操作,比你自己写for i in range(len(df))这种逐行循环效率高得多,尤其是数据量较大的时候。

内容的提问来源于stack exchange,提问作者PiotrK

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 10:44:40