使用imputeTS包中gplot_na_imputations()或ggplot_na_distribution()时遇输入非数值错误的解决方案咨询
Why the Error Happens
The ggplot_na_imputations() and ggplot_na_distribution() functions from imputeTS are built to work with univariate (single) numeric time series—think a vector, ts object, or a numeric dataframe with one column. Your current dataframe includes a non-numeric countries column, and each row represents a separate time series (one per country) spread across month columns. Passing the entire dataframe directly to these functions triggers the "not numeric" error because they can't handle the mix of character data and multi-column numeric time series.
On top of that, your current imputation step (na_kalman(total_tests_md)) is likely doing something unintended: it fills missing values column-wise (using other countries' data for the same month) instead of row-wise (using a country's own historical/future time points for imputation). Let's fix both issues.
Step 1: Correctly Impute Time Series Per Country
First, let's restructure your data to handle each country's time series properly:
Reshape to long (tidy) format: This makes it easy to group and process each country's data individually:
library(tidyverse) library(imputeTS) # Convert wide dataframe to long format total_tests_long <- total_tests_md %>% pivot_longer(cols = -countries, names_to = "month", values_to = "tests") %>% # Convert month strings to proper dates (for time ordering) mutate(month = lubridate::my(month)) %>% arrange(countries, month)Impute missing values for each country:
total_tests_imputed <- total_tests_long %>% group_by(countries) %>% mutate(tests_imputed = na_kalman(tests)) %>% ungroup()
Step 2: Plotting Solutions
Now you can generate plots for individual countries or batch-process all countries.
Option A: Plot a Single Country's Imputations
Pick a country (e.g., Albania) and extract its original and imputed time series:
# Extract Albania's original and imputed data albania_original <- total_tests_long %>% filter(countries == "Albania") %>% pull(tests) albania_imputed <- total_tests_imputed %>% filter(countries == "Albania") %>% pull(tests_imputed) # Convert to ts object (adds time series context like start date/frequency) albania_ts <- ts(albania_original, start = c(2020, 1), frequency = 12) albania_ts_imp <- ts(albania_imputed, start = c(2020, 1), frequency = 12) # Plot missing value distribution ggplot_na_distribution(albania_ts) # Plot imputations against original data ggplot_na_imputations(x_with_na = albania_ts, x_with_imputations = albania_ts_imp)
Option B: Generate Plots for All Countries (Loop)
Use a loop to create and display/save plots for every country:
# Get list of unique countries countries_list <- unique(total_tests_long$countries) # Loop through each country for (country in countries_list) { # Extract original and imputed time series original_data <- total_tests_long %>% filter(countries == country) %>% pull(tests) imputed_data <- total_tests_imputed %>% filter(countries == country) %>% pull(tests_imputed) # Convert to ts objects ts_original <- ts(original_data, start = c(2020, 1), frequency = 12) ts_imputed <- ts(imputed_data, start = c(2020, 1), frequency = 12) # Create and display plots dist_plot <- ggplot_na_distribution(ts_original) + ggtitle(paste("Missing Value Distribution:", country)) imp_plot <- ggplot_na_imputations(ts_original, ts_imputed) + ggtitle(paste("Imputed Values:", country)) print(dist_plot) print(imp_plot) # Optional: Save plots to files # ggsave(paste0("na_dist_", country, ".png"), dist_plot) # ggsave(paste0("imputations_", country, ".png"), imp_plot) }
Key Takeaways
- Always pass a single numeric vector or
tsobject to imputeTS plotting functions—exclude or split out non-numeric columns first. - Using long-format data simplifies per-country time series operations (imputation, plotting) with tidyverse tools.
- Converting to a
tsobject adds helpful time series metadata (like start year and monthly frequency) that improves plot clarity.
内容的提问来源于stack exchange,提问作者Fabricio Sousa

