咨询:基于Pima数据集实现Triceps与多因变量的简单线性回归绘图
Visualizing Linear Relationships Between Triceps and Key Variables in the Pima Dataset
It looks like you're aiming to create three separate scatter plots, each showing the linear relationship between triceps (on the x-axis) and one of pregnant, glucose, or pressure (on the y-axis) using the Pima dataset. Here's a complete, cleaned-up version of your code that finishes the first plot and adds the other two, plus some practical best practices:
# Load required libraries library(ggplot2) library(pdp) # Retained as per your original code, even if not used directly here # Prepare dataset by removing rows with missing values pima_clean <- na.omit(pima) head(pima_clean) # Plot 1: Triceps vs Number of Pregnancies lin1 <- lm(triceps ~ pregnant, data = pima_clean) coef1 <- coef(lin1) p1 <- ggplot(data = pima_clean, aes(x = pregnant, y = triceps)) + geom_point(alpha = 0.5) + # Add alpha to reduce overplotting geom_abline(color = "blue", intercept = coef1[1], slope = coef1[2], size = 1) + labs( title = "Triceps Skin Fold vs Number of Pregnancies", x = "Number of Pregnancies", y = "Triceps Skin Fold Thickness (mm)" ) + theme_minimal() # Plot 2: Triceps vs Glucose Level lin2 <- lm(triceps ~ glucose, data = pima_clean) coef2 <- coef(lin2) p2 <- ggplot(data = pima_clean, aes(x = glucose, y = triceps)) + geom_point(alpha = 0.5) + geom_abline(color = "red", intercept = coef2[1], slope = coef2[2], size = 1) + labs( title = "Triceps Skin Fold vs Plasma Glucose Concentration", x = "Glucose Level (mg/dL)", y = "Triceps Skin Fold Thickness (mm)" ) + theme_minimal() # Plot 3: Triceps vs Blood Pressure lin3 <- lm(triceps ~ pressure, data = pima_clean) coef3 <- coef(lin3) p3 <- ggplot(data = pima_clean, aes(x = pressure, y = triceps)) + geom_point(alpha = 0.5) + geom_abline(color = "green", intercept = coef3[1], slope = coef3[2], size = 1) + labs( title = "Triceps Skin Fold vs Diastolic Blood Pressure", x = "Blood Pressure (mm Hg)", y = "Triceps Skin Fold Thickness (mm)" ) + theme_minimal() # Optional: Arrange all plots in a grid for easy comparison # install.packages("gridExtra") if you don't have it library(gridExtra) grid.arrange(p1, p2, p3, ncol = 2)
Quick Tips:
- I renamed the cleaned dataset to
pima_cleanto avoid overwriting the raw original data—this is a safe practice for reproducible analysis. - Swapped
pima$columnfor unquoted column names insideaes()—this is the standard ggplot2 syntax and keeps code cleaner. - Added
alpha = 0.5to points to make overcrowded areas easier to interpret, showing data density instead of overlapping dots. - Included descriptive labels and titles to make your plots self-explanatory for anyone viewing them.
内容的提问来源于stack exchange,提问作者Ville
相关产品推荐
相关产品推荐

