如何用ggplot将多个函数整合到一张图表中?附数据集详情
Got it, let's break this down for you. You have a 25×6 dataframe with factor data (Mkt.RF, SMB, HML, RMW, CMA, WML), and want to combine multiple related plots or series into one visualization with ggplot2. Here's a step-by-step guide with concrete examples:
First: Prep Your Data (Critical!)
ggplot2 works best with long-format data (one row per observation per variable) instead of your current wide format (one column per factor). Let's start by converting it using the tidyr package:
First, let's simulate a dataframe matching your structure (since your raw data was cut off):
# Load required packages library(ggplot2) library(tidyr) library(dplyr) # Simulate your 25x6 dataframe (replace this with your actual data) set.seed(123) df <- data.frame( Mkt.RF = rnorm(25), SMB = rnorm(25), HML = rnorm(25), RMW = rnorm(25), CMA = rnorm(25), WML = rnorm(25) ) # Add a time/observation index (since you have 25 rows, assume this is time) df$obs <- 1:nrow(df)
Now convert to long format:
df_long <- df %>% pivot_longer(cols = -obs, names_to = "Factor", values_to = "Return")
This gives you a dataframe with 3 columns: obs (your x-axis, e.g., time), Factor (the variable name), and Return (the value).
Option 1: Overlay All Series on One Plot
If you want to compare all factors directly on the same axes, use color to distinguish them:
ggplot(df_long, aes(x = obs, y = Return, color = Factor)) + geom_line(linewidth = 1) + # Use lines for time series geom_point(size = 2) + # Add points for clarity labs( title = "Factor Returns Over Time", x = "Observation Number", y = "Return", color = "Factor" ) + theme_minimal() + scale_color_viridis_d(option = "plasma") # Nice color palette for categorical variables
This will show all 6 factor series on one plot, each with a unique color.
Option 2: Faceted Plots (Separate Panels for Each Factor)
If overlaying gets too cluttered, use facets to create a grid of small plots (one per factor):
ggplot(df_long, aes(x = obs, y = Return)) + geom_line(color = "#2c3e50", linewidth = 1) + geom_point(color = "#e74c3c", size = 2) + facet_wrap(~Factor, ncol = 2) + # Arrange facets in 2 columns labs( title = "Factor Returns by Factor", x = "Observation Number", y = "Return" ) + theme_minimal() + theme(strip.background = element_rect(fill = "#f8f9fa"), strip.text = element_text(face = "bold"))
This keeps each factor's data separate but aligned for easy comparison.
Bonus: Adding Derived Functions (e.g., Rolling Averages)
If you want to plot the original data alongside a derived function (like a rolling mean), calculate it first and merge it into your long data:
# Calculate rolling 3-period mean for each factor df_rolling <- df_long %>% group_by(Factor) %>% mutate(Rolling_Mean = zoo::rollmean(Return, k = 3, fill = NA)) # Use zoo package for rolling stats # Plot original returns + rolling mean ggplot(df_rolling, aes(x = obs, y = Return)) + geom_line(color = "gray50", alpha = 0.6) + geom_line(aes(y = Rolling_Mean), color = "#2980b9", linewidth = 1.2) + facet_wrap(~Factor, ncol = 2) + labs( title = "Factor Returns with 3-Period Rolling Mean", x = "Observation Number", y = "Value" ) + theme_minimal()
This adds a smoothed rolling average line on top of each factor's raw return data.
Just replace the simulated data with your actual dataframe, and adjust the x-axis (if your rows are dates instead of observation numbers, use as.Date() to format it properly).
内容的提问来源于stack exchange,提问作者Mads

