You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于R语言的数据模式可视化:附相关数值数据表

Visualizing Data Patterns in R for Your Numerical Dataset

Got it, let's walk through how to turn your raw numerical data into meaningful visualizations step by step. First, we need to structure your data (it looks like groups of three values per observation) then use R's plotting tools to spot patterns.

Step 1: Prepare the Data

First, let's get your data into a tidy data frame and load the tools we need:

# Install and load essential packages if missing
if (!require("tidyverse")) {
  install.packages("tidyverse")
  library(tidyverse)
}

# Your raw data (add the remaining values if you have the full set)
raw_data <- c(0.5625, 0.037037037, 0.009923785, 
              0.7734375, 0.0781893, 0.009923785, 
              0.9609375, 0.127572016, 0.009923785, 
              0.26953125, 0.008230453, 0.009923785, 
              0.85546875, 0.144032922, 0.009923785, 
              0.873046875, 0.187928669, 0.009923785, 
              0.969726563, 0.138545953, 0.009923785, 
              0.711914063, 0.031550069, 0.009923785, 
              0.588867188, 0.066300869, 0.009923785, 
              0.670898438, 0.038866027, 0.009923785, 
              0.331054688, 0.004572474, 0.009328358, 
              0.670898438, 0.038866027, 0.009923785, 
              0.8203125, 0.1015625, 0.009923785, 
              0.794921875, 0.115234375, 0.009923785)

# Convert to a structured data frame (3 columns per observation)
df <- matrix(raw_data, ncol = 3, byrow = TRUE) %>%
  as.data.frame() %>%
  rename(Var1 = V1, Var2 = V2, Var3 = V3) %>%
  mutate(Observation = row_number()) # Add an ID to track order

Step 2: Check Variable Distributions

First, let's see how each variable is spread out with boxplots—this helps spot outliers or consistent patterns:

df %>%
  pivot_longer(cols = starts_with("Var"), names_to = "Variable", values_to = "Value") %>%
  ggplot(aes(x = Variable, y = Value)) +
  geom_boxplot(fill = "lightblue", alpha = 0.7) +
  labs(title = "Distribution of Each Variable", x = "Variable", y = "Value") +
  theme_minimal()

You’ll notice Var3 is almost entirely constant (~0.0099) with just one tiny outlier, while Var1 and Var2 have much more variability.

Step 3: Explore Relationships Between Variables

Use scatter plots to see if variables move together:

# Pairwise scatter plot matrix
pairs(df[,1:3], 
      main = "How Variables Relate to Each Other",
      pch = 16, 
      col = "darkgreen")

# Or a polished ggplot version for Var1 vs Var2
df %>%
  ggplot(aes(x = Var1, y = Var2)) +
  geom_point(color = "darkred", size = 2) +
  labs(title = "Var1 vs Var2", x = "Var1", y = "Var2") +
  theme_minimal()

This will show you a clear positive correlation: when Var1 increases, Var2 tends to increase too. Var3 doesn’t correlate with either, which makes sense given its consistency.

If your data is ordered (like time series), plot how each variable changes across observations:

df %>%
  pivot_longer(cols = starts_with("Var"), names_to = "Variable", values_to = "Value") %>%
  ggplot(aes(x = Observation, y = Value, color = Variable)) +
  geom_line(size = 1) +
  geom_point(size = 1.5) +
  labs(title = "Variable Trends Across Observations", x = "Observation", y = "Value") +
  theme_minimal() +
  scale_color_manual(values = c("Var1" = "darkorange", "Var2" = "purple", "Var3" = "darkgreen"))

Here you’ll see Var3 stays flat, while Var1 and Var2 fluctuate in sync—their peaks and valleys line up nicely.

Step 5: Heatmap for Value Magnitude

A heatmap makes it easy to spot which observations have high/low values for each variable:

df %>%
  select(-Observation) %>%
  t() %>%
  as.data.frame() %>%
  rownames_to_column("Variable") %>%
  pivot_longer(cols = -Variable, names_to = "Observation", values_to = "Value") %>%
  mutate(Observation = as.numeric(str_remove(Observation, "V"))) %>%
  ggplot(aes(x = Observation, y = Variable, fill = Value)) +
  geom_tile() +
  scale_fill_viridis_c(option = "plasma") +
  labs(title = "Heatmap of Variable Values", x = "Observation", y = "Variable") +
  theme_minimal()

Dark colors represent higher values, so you’ll quickly see which observations have the highest Var1 and Var2 scores.

Key Patterns to Note

From these plots, it’s clear:

  • Var3 is effectively a constant variable (only one minor deviation)
  • Var1 and Var2 have a strong positive linear relationship
  • Fluctuations in Var1 directly mirror fluctuations in Var2

Feel free to tweak colors, labels, or add more observations to the raw data if you have the full dataset!

内容的提问来源于stack exchange,提问作者vaibhav Mishra

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:43:04