You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

在Google Colab中用Geopandas绘制伊朗事故Shapefile时遇报错求助

解决Geopandas可视化伊朗事故Shapefile时的NaN转整数报错问题

问题场景

在Google Colab中读取伊朗事故Shapefile为Geopandas DataFrame后,调用plot函数时触发以下错误:

ValueError: cannot convert float NaN to integer

同时伴随警告:

/usr/local/lib/python3.7/dist-packages/matplotlib/axes/_base.py:2450: RuntimeWarning: overflow encountered in double_scalars
WARNING:matplotlib.text:posx and posy should be finite values

用户使用的代码:

import pandas as pd
import matplotlib.pyplot as plt
import geopandas as gpd
gdf = gpd.read_file('/content/accidents.shp',crs=4326)

gdf = gdf.fillna(value=0)
fig, ax = plt.subplots(1, figsize=(20, 20))
ax.axis('off')
ax.set_title('accidents in Iran',
             fontdict={'fontsize': '15', 'fontweight' : '3'})
fig = gdf.plot(column='tedad_jarh', cmap='RdYlGn', linewidth=0.5, ax=ax, edgecolor='0.2',legend=True)

输入数据说明:数据包含tedad_jarh(伤亡类数值字段)和多边形几何列,已用fillna(0)填充空值。


解决建议

1. 修正坐标参考系(CRS)设置

读取Shapefile时强行指定crs=4326可能覆盖文件自带的正确CRS,导致坐标异常引发计算溢出:

  • 先读取文件并查看原CRS:
    gdf = gpd.read_file('/content/accidents.shp')
    print("原文件CRS:", gdf.crs)
    
  • 若原CRS不是4326,用to_crs转换而非读取时指定:
    gdf = gdf.to_crs(epsg=4326)
    
  • 检查并修复无效几何:
    # 查看无效几何数量
    print(gdf.geometry.is_valid.value_counts())
    # 修复无效多边形(通用方法)
    gdf.geometry = gdf.geometry.buffer(0)
    

2. 排查tedad_jarh列数据问题

即使做了fillna(0),仍可能存在非数值类型或极端值导致计算错误:

  • 检查字段类型和数据范围:
    print("字段类型:", gdf['tedad_jarh'].dtype)
    print("数据统计:\n", gdf['tedad_jarh'].describe())
    
  • 强制转换为数值类型并填充异常值:
    gdf['tedad_jarh'] = pd.to_numeric(gdf['tedad_jarh'], errors='coerce').fillna(0)
    

3. 调整图例参数避免NaN计算

报错源于图例刻度的NaN计算,可手动控制数值范围或切换图例类型:

  • 手动指定图例的数值范围:
    fig, ax = plt.subplots(1, figsize=(20, 20))
    ax.axis('off')
    ax.set_title('accidents in Iran', fontdict={'fontsize': '15', 'fontweight' : '3'})
    # 用字段的最小/最大值约束图例范围
    gdf.plot(column='tedad_jarh', cmap='RdYlGn', linewidth=0.5, ax=ax, edgecolor='0.2',
             legend=True, vmin=gdf['tedad_jarh'].min(), vmax=gdf['tedad_jarh'].max())
    
  • 改用分箱分类图例(避免连续数值的浮点计算问题):
    # 对tedad_jarh做分箱处理
    gdf['jarh_bins'] = pd.cut(gdf['tedad_jarh'], bins=5, labels=['极低', '低', '中等', '高', '极高'])
    fig, ax = plt.subplots(1, figsize=(20, 20))
    ax.axis('off')
    ax.set_title('accidents in Iran', fontdict={'fontsize': '15', 'fontweight' : '3'})
    gdf.plot(column='jarh_bins', cmap='RdYlGn', linewidth=0.5, ax=ax, edgecolor='0.2', legend=True)
    

4. 逐步简化代码定位问题

先排除基础绘图问题,再逐步添加参数:

  • 先绘制无样式的基础地图:
    fig, ax = plt.subplots(1, figsize=(20,20))
    gdf.plot(ax=ax)
    plt.show()
    
  • 若基础地图正常,再依次添加column、cmap、legend等参数,定位具体引发错误的设置。

内容的提问来源于stack exchange,提问作者fatane

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 00:10:36