如何将含WK_CODE字段的GeoJSON文件与Pandas DataFrame合并并按WK_CODE统计人口
Let’s walk through how to merge your GeoJSON file with the CSV-derived DataFrame, then compute population statistics grouped by WK_CODE:
1. Install Required Libraries
First, make sure you have the necessary tools installed (if you don’t already):
pip install geopandas pandas
2. Load Your Data Files
Use geopandas to handle the GeoJSON (it preserves both spatial and attribute data) and pandas for the CSV:
import geopandas as gpd import pandas as pd # Load GeoJSON into a GeoDataFrame geo_data = gpd.read_file("your_geojson_file.geojson") # Load CSV into a regular Pandas DataFrame csv_data = pd.read_csv("your_csv_file.csv")
3. Fix Data Type Mismatches for WK_CODE
Merges can fail if WK_CODE is stored as different types (e.g., string vs integer) in the two datasets. Let’s standardize it to strings to avoid issues:
geo_data["WK_CODE"] = geo_data["WK_CODE"].astype(str) csv_data["WK_CODE"] = csv_data["WK_CODE"].astype(str)
4. Merge the Datasets
Combine the two datasets using WK_CODE as the key. Choose a join type based on your needs:
inner: Keep only rows with matchingWK_CODEin both files (default)left: Keep all entries from the GeoJSON, even if there’s no CSV matchright: Keep all entries from the CSV, even if there’s no GeoJSON match
# Example: Inner merge to keep only matching records merged_data = pd.merge(geo_data, csv_data, on="WK_CODE", how="inner") # If you want to keep spatial geometry as a GeoDataFrame: merged_geo_data = geo_data.merge(csv_data, on="WK_CODE", how="inner")
5. Calculate Population Statistics by WK_CODE
Now group the merged data by WK_CODE and compute your desired stats (sum, average, count, etc.). Note: Your sample GeoJSON has POPULATION stored as a string—convert it to numeric first for calculations:
# Convert population column to numeric (handle any invalid values gracefully) merged_data["POPULATION"] = pd.to_numeric(merged_data["POPULATION"], errors="coerce") # Group by WK_CODE and calculate stats population_summary = merged_data.groupby("WK_CODE").agg( total_population=("POPULATION", "sum"), average_population=("POPULATION", "mean"), record_count=("POPULATION", "count") ).reset_index() # Print or save the result print(population_summary) # population_summary.to_csv("population_stats.csv", index=False)
If you want to keep the spatial geometry alongside stats (for mapping later), group by both WK_CODE and geometry:
geo_pop_summary = merged_geo_data.groupby(["WK_CODE", "geometry"]).agg( total_population=("POPULATION", "sum") ).reset_index()
Quick Notes
- Replace
"your_geojson_file.geojson"and"your_csv_file.csv"with your actual file paths. - Adjust the aggregation functions (
sum,mean, etc.) to match the stats you need.
内容的提问来源于stack exchange,提问作者Antonius K

