如何解读R语言plot函数生成绘图中的竖线?相关疑问咨询
Hey there! Let's break down what's happening with your plot and how to adjust it to get the exact values you're looking for:
First, let's unpack those confusing vertical lines:
You used
type = "l"to draw a line plot, but yourmedicare$ageyrslikely has multipleothervalues (hospitalization counts per 10k people) tied to the same age. When R renders a line plot with duplicate x-values, it connects all corresponding y-values in the order they appear in your dataset—this creates vertical lines that represent the full range ofothervalues for each age group. The vertical line at age 65 specifically shows the min-to-max spread of hospitalization counts for that age bracket.
How to Plot Exact Single Values Instead of Ranges
To visualize a single representative value per age (like a mean or median) instead of the full range, you'll need to summarize your data first. Here are two straightforward ways to do this:
Option 1: Base R Summary
# Calculate mean of `other` for each age, ignoring NA values age_means <- tapply(medicare$other, medicare$ageyrs, mean, na.rm = TRUE) # Convert to a data frame for easier plotting summary_df <- data.frame(ageyrs = as.numeric(names(age_means)), other_mean = age_means) # Plot the summarized data plot(summary_df$ageyrs, summary_df$other_mean, type = "l", xlab = "Age", ylab = "Mean Hospitalizations per 10k", main = "Figure 2: Mean Hospitalizations by Age")
Option 2: Using dplyr (More Readable for Complex Summaries)
library(dplyr) # Create a summary data frame with mean and median per age medicare_summary <- medicare %>% group_by(ageyrs) %>% summarise( other_mean = mean(other, na.rm = TRUE), other_median = median(other, na.rm = TRUE) ) # Plot the median (replace with other_mean if you prefer) plot(medicare_summary$ageyrs, medicare_summary$other_median, type = "l", xlab = "Age", ylab = "Median Hospitalizations per 10k", main = "Figure 2: Median Hospitalizations by Age")
Clarifying the "Near-Horizontal" Lines
Your hunch is correct! If you see flat segments between age groups in the summarized plot, that does represent the change (or lack of change) in your other metric across age groups. In the raw data plot, those flat-looking areas might be an illusion caused by overlapping ranges between adjacent ages—but using summarized single values will show you the true trend clearly.
Bonus: Visualizing Ranges Without Confusing Vertical Lines
If you still want to display the range of values but avoid messy vertical lines, try a boxplot or error bar plot instead:
# Boxplot to show full distribution per age boxplot(other ~ ageyrs, data = medicare, xlab = "Age", ylab = "Hospitalizations per 10k", main = "Figure 2: Distribution of Hospitalizations by Age") # Error bar plot showing mean ± standard deviation medicare_summary <- medicare %>% group_by(ageyrs) %>% summarise( mean_val = mean(other, na.rm = TRUE), sd_val = sd(other, na.rm = TRUE) ) plot(medicare_summary$ageyrs, medicare_summary$mean_val, type = "p", pch = 16, xlab = "Age", ylab = "Hospitalizations per 10k", main = "Figure 2: Mean Hospitalizations with SD Error Bars") # Add error bars arrows(medicare_summary$ageyrs, medicare_summary$mean_val - medicare_summary$sd_val, medicare_summary$ageyrs, medicare_summary$mean_val + medicare_summary$sd_val, length = 0.05, angle = 90, code = 3)
内容的提问来源于stack exchange,提问作者Collective Action

