You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中优化大数据集散点图?含DATE/INT_RATE按GRADE过滤绘图需求

Hey there! Let's walk through how to create your desired plot with GRADE filtering, then cover some key tricks to make scatterplots work better for large datasets in R.

1. First: Clean Your Data & Plot with GRADE Filtering

First off, your raw data needs a quick cleanup: INT_RATE is a string with a % sign, and DATE is in a shorthand format—we'll fix those first. Here's how to do it with the tidyverse (ggplot2 + dplyr):

library(tidyverse)

# Assume your dataset is named `raw_df`
clean_df <- raw_df %>%
  # Convert interest rate to numeric (strip the % and turn to number)
  mutate(INT_RATE = as.numeric(str_remove(INT_RATE, "%")),
         # Convert date to a proper date type (add a year, e.g., 2023, adjust as needed)
         DATE = dmy(paste0(DATE, "-2023")))

# Plot only Grade A data
ggplot(clean_df %>% filter(GRADE == "A"), aes(x = DATE, y = INT_RATE)) +
  geom_point() +
  labs(title = "Date vs. Interest Rate (Grade A Only)",
       x = "Date", y = "Interest Rate (%)") +
  theme_minimal()

# If you want to compare all grades at once, use color to distinguish:
ggplot(clean_df, aes(x = DATE, y = INT_RATE, color = GRADE)) +
  geom_point() +
  labs(title = "Date vs. Interest Rate by Grade",
       x = "Date", y = "Interest Rate (%)") +
  theme_minimal()

2. Optimizing Scatterplots for Large Datasets

When dealing with massive datasets, regular scatterplots suffer from overplotting—points stack up so much you can't see patterns. Here are the most effective fixes:

Add Transparency (Alpha)

Lower the opacity of points so overlapping areas appear darker, revealing density patterns:

ggplot(clean_df, aes(x = DATE, y = INT_RATE, color = GRADE)) +
  geom_point(alpha = 0.1) # Adjust alpha between 0 (fully transparent) and 1 (opaque)

Random Sampling

If your dataset is way too big, sample a subset—you'll retain the overall trend without sacrificing speed:

# Sample 10% of the data (adjust the fraction as needed)
sampled_df <- clean_df %>% sample_frac(0.1)

ggplot(sampled_df, aes(x = DATE, y = INT_RATE, color = GRADE)) +
  geom_point()

Hexbin or 2D Binning

Instead of individual points, use hexagons (or rectangles) to show how many points fall into each area—perfect for extremely large datasets:

# Hexbin plot (shows density with color)
ggplot(clean_df, aes(x = DATE, y = INT_RATE)) +
  geom_hex() +
  scale_fill_viridis_c(name = "Number of Points") +
  labs(title = "Date vs. Interest Rate Density",
       x = "Date", y = "Interest Rate (%)")

# 2D histogram (rectangular bins)
ggplot(clean_df, aes(x = DATE, y = INT_RATE)) +
  geom_bin2d() +
  scale_fill_viridis_c(name = "Count")

Interactive Plots

Use plotly to make your scatterplot interactive—you can zoom into specific regions, hover to see individual data points, and pan around:

library(plotly)

base_plot <- ggplot(clean_df, aes(x = DATE, y = INT_RATE, color = GRADE)) +
  geom_point(alpha = 0.2) +
  theme_minimal()

# Convert to interactive plot
ggplotly(base_plot)

Aggregate Data

Instead of plotting every single point, aggregate by date and grade to show trends (like average interest rate per date):

aggregated_df <- clean_df %>%
  group_by(DATE, GRADE) %>%
  summarise(Average_Rate = mean(INT_RATE, na.rm = TRUE)) %>%
  ungroup()

ggplot(aggregated_df, aes(x = DATE, y = Average_Rate, color = GRADE)) +
  geom_line(linewidth = 1) +
  labs(title = "Average Interest Rate Over Time by Grade",
       x = "Date", y = "Average Interest Rate (%)") +
  theme_minimal()

内容的提问来源于stack exchange,提问作者scugn1zz0

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 06:51:04