R语言中基于事故严重度加权的2D Kernel Density实现咨询
Great question! Let's walk through how to handle weighted 2D kernel density estimation in R for your road accident data, especially with your severity-based weighting scheme (3x for fatal accidents, 1x for property-only incidents). We'll cover two practical approaches, depending on your data size and needs.
Approach 1: Duplicate Observations (Simple & Intuitive)
If your dataset isn't extremely large, this method is the easiest to grasp. It directly simulates your weighting rule by repeating each observation according to its severity weight:
Define raw weights based on accident severity:
# Assume your dataset is named `accident_data` with columns: lon, lat, severity accident_data$weight_raw <- ifelse(accident_data$severity == "Fatal", 3, 1)Create a weighted dataset by repeating rows:
# Generate indices to repeat rows according to their weight repeat_indices <- rep(1:nrow(accident_data), times = accident_data$weight_raw) weighted_data <- accident_data[repeat_indices, c("lon", "lat")]Compute 2D kernel density using the
MASSpackage'skde2d()function:library(MASS) # `n` controls the resolution of the density grid (adjust as needed) kde_result <- kde2d(weighted_data$lon, weighted_data$lat, n = 100)Visualize the result:
image(kde_result, main = "Weighted 2D Kernel Density of Road Accidents") points(accident_data$lon, accident_data$lat, pch = 16, cex = 0.5, col = "red")
This approach avoids any weight scaling headaches, but note that it can bloat your dataset (e.g., 1000 fatal accidents become 3000 rows), which might be an issue for very large datasets.
Approach 2: Use a Weight-Aware Kernel Density Function (Efficient for Large Data)
For larger datasets, use the ks package, which natively supports weighted kernel density estimation without needing to duplicate rows. It also doesn't require weights to sum to the sample size (unlike stats::density() for 1D cases):
Install and load the
kspackage:install.packages("ks") library(ks)Define your raw weights (no scaling needed):
accident_data$weight <- ifelse(accident_data$severity == "Fatal", 3, 1)Compute weighted 2D kernel density:
# Convert coordinates to a matrix (required by ks::kde()) coords <- as.matrix(accident_data[, c("lon", "lat")]) # Pass weights directly to the `w` parameter kde_weighted <- kde(x = coords, w = accident_data$weight)Visualize the result:
plot(kde_weighted, main = "Weighted 2D Kernel Density (ks Package)") points(coords, pch = 16, cex = 0.5, col = "blue")
Bonus: Handling 1D Weighted Density (For Reference)
If you ever need to do 1D weighted density (e.g., just longitude or latitude) with stats::density(), you'll need to scale weights so their sum equals the sample size, as you noted:
weight_raw <- ifelse(accident_data$severity == "Fatal", 3, 1) n <- length(weight_raw) # Scale weights to sum to sample size weight_scaled <- weight_raw * (n / sum(weight_raw)) # Compute weighted 1D density density_1d <- density(accident_data$lon, w = weight_scaled) plot(density_1d, main = "Weighted 1D Kernel Density of Accident Longitudes")
内容的提问来源于stack exchange,提问作者NOA11120

