You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在R中绘制随时间变化的相对比例图?含研究绘图需求

Got it, let's break down how to create this exact chart you need for your genre annotation study. I’ve worked through similar visualization tasks for behavioral experiments, so here’s a practical, step-by-step approach with code examples to get you sorted:

Core Requirements Recap

First, let's align on the key details to make sure we're building the right thing:

  • X-axis: Participants ordered chronologically, with all tags from the same participant sharing the exact same X position (e.g., Tagid 1 and 2 from Participant A both map to X=3 if they were the 3rd participant)
  • Y-axis: Frequency of your target tag word (either raw count of how many times the participant used the word, or its proportion relative to all tags they submitted)
  • Dataset: 10 songs' genre annotation data, with each participant's tags stored separately
Step 1: Data Preprocessing

Before plotting, you’ll need to clean and restructure your data into a plot-friendly format:

  • Assign a unique chronological participant ID (e.g., Participant_1, Participant_2) to every tag entry, matching the order participants completed the task
  • Add a time order index (e.g., 1, 2, 3) for each participant — this will be your X-axis value (same index for all tags from one participant)
  • Calculate the target word's frequency per participant (raw count or proportion)
Step 2: Python Implementation (Matplotlib/Seaborn)

This is a go-to setup for scientific plotting. Below is a complete example with simulated data matching your use case:

import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns

# Simulate your actual dataset structure
raw_data = {
    "Participant_ID": ["P1", "P1", "P2", "P3", "P3", "P3", "P4"],
    "Tagid": [1, 2, 1, 1, 2, 3, 1],
    "Tag_Word": ["jazz", "blues", "jazz", "rock", "jazz", "jazz", "blues"],
    "Time_Order": [1, 1, 2, 3, 3, 3, 4]  # X-axis position for each participant
}
df = pd.DataFrame(raw_data)

# Define your target tag word
target_tag = "jazz"

# Calculate raw frequency (count of target word per participant)
frequency_df = df[df["Tag_Word"] == target_tag].groupby("Participant_ID")["Tagid"].count().reset_index(name="Target_Count")
# Merge with time order data (ensure one X value per participant)
frequency_df = frequency_df.merge(df[["Participant_ID", "Time_Order"]].drop_duplicates(), on="Participant_ID")

# Plot the chart
plt.figure(figsize=(12, 6))

# Bar chart for overall participant frequency
sns.barplot(x="Time_Order", y="Target_Count", data=frequency_df, color="#2ca02c", alpha=0.7)

# Optional: Add scatter points to show individual tag occurrences (same X for same participant)
individual_tags = df[df["Tag_Word"] == target_tag]
plt.scatter(x=individual_tags["Time_Order"], y=[0.3]*len(individual_tags), color="#ff7f0e", s=50, label="Individual Tag Use")

# Customize labels and styling
plt.xlabel("Participants (Chronological Order)", fontsize=12)
plt.ylabel(f"Frequency of '{target_tag}'", fontsize=12)
plt.title(f"'{target_tag}' Usage Across 10-Song Genre Annotation Participants", fontsize=14)
plt.xticks(rotation=45)
plt.legend()
plt.tight_layout()
plt.show()
Step 3: Adjustments for Proportion-Based Frequency

If you want Y-axis to show the proportion of the target word relative to all tags a participant used, modify the frequency calculation like this:

# Calculate total tags per participant
total_tags = df.groupby("Participant_ID")["Tagid"].count().reset_index(name="Total_Tags")
# Calculate target word count per participant
target_counts = df[df["Tag_Word"] == target_tag].groupby("Participant_ID")["Tagid"].count().reset_index(name="Target_Count")
# Merge and compute proportion
frequency_df = pd.merge(total_tags, target_counts, on="Participant_ID", how="left").fillna(0)
frequency_df["Target_Proportion"] = frequency_df["Target_Count"] / frequency_df["Total_Tags"]
# Merge with time order data
frequency_df = frequency_df.merge(df[["Participant_ID", "Time_Order"]].drop_duplicates(), on="Participant_ID")

Then use Target_Proportion as your Y-axis value in the plot.

Alternative: R Implementation (ggplot2)

If you prefer R for your analysis, here's an equivalent setup:

library(ggplot2)
library(dplyr)

# Simulate data
raw_data <- data.frame(
  Participant_ID = c("P1", "P1", "P2", "P3", "P3", "P3", "P4"),
  Tagid = c(1, 2, 1, 1, 2, 3, 1),
  Tag_Word = c("jazz", "blues", "jazz", "rock", "jazz", "jazz", "blues"),
  Time_Order = c(1, 1, 2, 3, 3, 3, 4)
)

target_tag <- "jazz"

# Calculate frequency
frequency_df <- raw_data %>%
  filter(Tag_Word == target_tag) %>%
  group_by(Participant_ID) %>%
  summarise(Target_Count = n()) %>%
  left_join(raw_data %>% distinct(Participant_ID, Time_Order), by = "Participant_ID")

# Plot
ggplot(frequency_df, aes(x = factor(Time_Order), y = Target_Count)) +
  geom_bar(stat = "identity", fill = "#2ca02c", alpha = 0.7) +
  geom_jitter(data = raw_data %>% filter(Tag_Word == target_tag), 
              aes(x = factor(Time_Order), y = 0.3), 
              color = "#ff7f0e", size = 3, alpha = 0.7) +
  labs(x = "Participants (Chronological Order)",
       y = paste("Frequency of '", target_tag, "'", sep = ""),
       title = "'jazz' Usage Across 10-Song Genre Annotation Participants") +
  theme(axis.text.x = element_text(rotation = 45, hjust = 1))

内容的提问来源于stack exchange,提问作者jbouwman

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 07:21:27