在R中创建印度分级统计图失败,求问题排查及解决方法
Hey Ravi, let's walk through the common issues with your current workflow and fix your choropleth map step by step:
1. Outdated Geospatial Tooling + Name Matching Issues
First off, rgdal::readOGR() and fortify() are legacy tools—rgdal is actually deprecated on CRAN now, and fortify often causes ID mismatches or broken polygon grouping. Worse, your state name fields (ST_NM from the shapefile vs Name in states_data) might have subtle differences (like whitespace, case, or special character formatting) that are breaking the merge, leaving gaps in your map.
Fix: Switch to the modern sf package
This package is designed for tidy geospatial work and avoids most of fortify's headaches:
# Install if you haven't already install.packages("sf") library(sf) library(dplyr) # Read the shapefile (sf automatically picks up all related .shp/.dbf/.prj files) india_states <- st_read("/Data/Admin2.shp") # Standardize names to eliminate matching errors india_states <- india_states %>% mutate(ST_NM = toupper(trimws(ST_NM))) states_data <- states_data %>% mutate(Name = toupper(trimws(Name))) # Merge directly (no fortify needed!) final_data <- merge(india_states, states_data, by.x = "ST_NM", by.y = "Name")
2. Wrong Grouping + Messed-Up Vertex Order
Your geom_polygon uses group=id, which is incorrect! When you fortify a shapefile, the group field represents individual polygon segments (critical for states with enclaves or multiple landmasses), while id is just the state name. Also, merge() shuffles the vertex order of your polygons, leading to distorted or overlapping shapes.
Fix: Use geom_sf (simpler and more reliable)
The sf package integrates directly with ggplot2, so you don't have to worry about grouping or vertex order:
library(ggplot2) ggplot(final_data) + # geom_sf handles all polygon logic automatically geom_sf(aes(fill = TOT_P), color = "black", size = 0.25) + # Use a projection optimized for India (EPSG:4326 is WGS84, universal for global maps) coord_sf(crs = st_crs(4326)) + # Add a readable color scale scale_fill_viridis_c(option = "plasma", name = "Total Population") + # Clean up the theme labs(title = "India State-Wise Population Choropleth") + theme_minimal()
If you absolutely need to stick with the fortify workflow, fix the grouping and sort vertices:
# Merge and re-sort by group to preserve polygon order final_data <- final_data %>% arrange(group) ggplot() + geom_polygon(data = final_data, aes(x = long, y = lat, group = group, fill = TOT_P), color = "black", size = 0.25) + coord_map(projection = "mercator")
3. Poor Projection Causing Distortion
coord_map() uses a default projection that stretches India's north-south axis. For a more accurate map, use a projection tailored to South Asia:
- EPSG:4326: WGS84 (universal, great for global context)
- EPSG:32643: UTM Zone 43N (covers northern India)
- EPSG:4275: India's official national coordinate system
4. Check for Unmatched Data
Always verify that your merge didn't drop any states or data points:
# Find states in the shapefile with no matching population data unmatched_states <- anti_join(india_states, states_data, by.x = "ST_NM", by.y = "Name") print(unmatched_states$ST_NM) # Find population entries with no matching state in the shapefile unmatched_data <- anti_join(states_data, india_states, by.x = "Name", by.y = "ST_NM") print(unmatched_data$Name)
Fix any name discrepancies (e.g., "JAMMU & KASHMIR" vs "JAMMU AND KASHMIR") manually to ensure full coverage.
内容的提问来源于stack exchange,提问作者Ravi

