如何用Python H3库获取美国区域resolution 11的H3 Index经纬度全表?
获取美国范围H3分辨率11索引及关联客户数据方案
核心思路
H3官方库本身没有直接提供「美国范围」的索引生成接口,需要先获取美国的地理边界,再用polyfill方法生成对应分辨率的H3索引,后续完成入库和关联操作。
步骤1:准备美国边界地理数据
先获取美国本土+飞地(阿拉斯加、夏威夷等)的GeoJSON格式边界文件,保存到本地(比如命名为usa_boundary.geojson)。
步骤2:用H3 Python库生成目标索引及经纬度
依赖安装
pip install h3-py geopandas pandas
代码实现
import h3 import geopandas as gpd import pandas as pd # 读取本地美国边界GeoJSON usa_gdf = gpd.read_file('usa_boundary.geojson') # 初始化集合存储H3索引(去重,避免跨多边形重复生成) h3_indices = set() # 遍历所有多边形(处理MultiPolygon类型的飞地) for geom in usa_gdf.geometry: # 转换为H3要求的GeoJSON坐标格式:[[[lon, lat], ...]] if geom.geom_type == 'MultiPolygon': for poly in geom.geoms: coords = [[list(coord) for coord in poly.exterior.coords]] # 生成分辨率11的H3索引,geo_json_conformant=True适配GeoJSON坐标顺序 batch_indices = h3.polyfill(coords, 11, geo_json_conformant=True) h3_indices.update(batch_indices) else: coords = [[list(coord) for coord in geom.exterior.coords]] batch_indices = h3.polyfill(coords, 11, geo_json_conformant=True) h3_indices.update(batch_indices) # 转换为带经纬度的DataFrame(H3索引对应中心点坐标) h3_records = [] for idx in h3_indices: lat, lon = h3.h3_to_geo(idx) # 返回(lat, lon) h3_records.append({ 'h3_index': idx, 'latitude': round(lat, 8), 'longitude': round(lon, 8) }) h3_df = pd.DataFrame(h3_records)
注意:分辨率11的H3索引在美国范围总数约千万级,生成时确保机器有足够内存,可分区域分批生成后合并。
步骤3:导入公司数据库
以PostgreSQL为例(其他数据库逻辑类似):
1. 数据库建表
CREATE TABLE h3_usa_res11 ( h3_index VARCHAR(20) PRIMARY KEY, latitude NUMERIC(10,8) NOT NULL, longitude NUMERIC(11,8) NOT NULL );
2. Python批量导入
from sqlalchemy import create_engine # 替换为你的数据库连接串 engine = create_engine('postgresql://用户名:密码@主机:端口/数据库名') # 导入数据,if_exists可选'replace'/'append' h3_df.to_sql('h3_usa_res11', engine, if_exists='replace', index=False)
步骤4:关联客户记录
方案1:数据库端关联(推荐,效率更高)
如果数据库支持H3扩展(如PostgreSQL的h3插件),直接通过客户经纬度生成H3索引后关联:
SELECT c.customer_id, c.customer_name, c.customer_lat, c.customer_lon, h.h3_index FROM customers c LEFT JOIN h3_usa_res11 h ON h.h3_index = h3_geo_to_h3(c.customer_lat, c.customer_lon, 11);
方案2:Python端预处理后关联
# 读取客户数据 customers_df = pd.read_sql('SELECT * FROM customers', engine) # 生成客户对应的H3分辨率11索引 customers_df['h3_index'] = customers_df.apply( lambda row: h3.geo_to_h3(row['customer_lat'], row['customer_lon'], 11), axis=1 ) # 关联H3数据表 merged_df = customers_df.merge(h3_df, on='h3_index', how='left') # 保存关联结果回数据库 merged_df.to_sql('customers_with_h3', engine, if_exists='replace', index=False)
注意事项
- 若美国边界GeoJSON文件过大,可先通过
geopandas.simplify()简化多边形,减少计算量 - 生成H3索引时可加入进度条(如
tqdm)监控进度 - 千万级数据入库时,建议分批次提交,避免数据库超时
内容的提问来源于stack exchange,提问作者Lily
相关产品推荐
相关产品推荐

