如何从ggplot()中获取与原始结构一致的处理后数据集?
Got it, let's work through how to get that filtered wide-format dataset you need directly from your ggplot object. Here's a straightforward, step-by-step approach:
Step 1: Extract ggplot's cleaned long-format data
First, make sure you've saved your ggplot object (let's call it p), then pull out the processed long-format data that ggplot actually used to generate your plot:
# Save your ggplot object first p <- ggplot(your_processed_long_data, aes(x = Time, y = value)) + geom_line() + facet_wrap(~name) # Extract the filtered long data ggplot used filtered_long_data <- ggplot_build(p)$plot$data
Step 2: Convert back to your original wide format
Now we just need to reshape this cleaned long data back into the Time + A1-A40 structure that matches your original dataset. Use pivot_wider to do this:
filtered_wide_data <- filtered_long_data %>% # Keep only the columns needed for reshaping select(Time, name, value) %>% # Reshape back to wide format (matches original Time + A1-A40 structure) pivot_wider(names_from = name, values_from = value) %>% # Optional: Reorder rows to match your original dataset's Time sequence arrange(Time)
Why this works
When ggplot drops those 33 rows, it's filtering out entries in the long-format data that have missing values (NA) or invalid points for plotting. By extracting this cleaned long data and pivoting it back to wide format, you'll get exactly the subset of your original wide data that ggplot used—no manual filtering required.
Quick note
If your original wide data had rows where all A1-A40 values were NA, ggplot would exclude those entire rows (which explains the 33 missing rows). The resulting filtered_wide_data will omit these rows, matching exactly what ggplot used for your plot.
内容的提问来源于stack exchange,提问作者Sophie Allan

