如何在R中绘制随时间变化的相对比例图?含研究绘图需求
Got it, let's break down how to create this exact chart you need for your genre annotation study. I’ve worked through similar visualization tasks for behavioral experiments, so here’s a practical, step-by-step approach with code examples to get you sorted:
First, let's align on the key details to make sure we're building the right thing:
- X-axis: Participants ordered chronologically, with all tags from the same participant sharing the exact same X position (e.g., Tagid 1 and 2 from Participant A both map to X=3 if they were the 3rd participant)
- Y-axis: Frequency of your target tag word (either raw count of how many times the participant used the word, or its proportion relative to all tags they submitted)
- Dataset: 10 songs' genre annotation data, with each participant's tags stored separately
Before plotting, you’ll need to clean and restructure your data into a plot-friendly format:
- Assign a unique chronological participant ID (e.g.,
Participant_1,Participant_2) to every tag entry, matching the order participants completed the task - Add a time order index (e.g.,
1,2,3) for each participant — this will be your X-axis value (same index for all tags from one participant) - Calculate the target word's frequency per participant (raw count or proportion)
This is a go-to setup for scientific plotting. Below is a complete example with simulated data matching your use case:
import pandas as pd import matplotlib.pyplot as plt import seaborn as sns # Simulate your actual dataset structure raw_data = { "Participant_ID": ["P1", "P1", "P2", "P3", "P3", "P3", "P4"], "Tagid": [1, 2, 1, 1, 2, 3, 1], "Tag_Word": ["jazz", "blues", "jazz", "rock", "jazz", "jazz", "blues"], "Time_Order": [1, 1, 2, 3, 3, 3, 4] # X-axis position for each participant } df = pd.DataFrame(raw_data) # Define your target tag word target_tag = "jazz" # Calculate raw frequency (count of target word per participant) frequency_df = df[df["Tag_Word"] == target_tag].groupby("Participant_ID")["Tagid"].count().reset_index(name="Target_Count") # Merge with time order data (ensure one X value per participant) frequency_df = frequency_df.merge(df[["Participant_ID", "Time_Order"]].drop_duplicates(), on="Participant_ID") # Plot the chart plt.figure(figsize=(12, 6)) # Bar chart for overall participant frequency sns.barplot(x="Time_Order", y="Target_Count", data=frequency_df, color="#2ca02c", alpha=0.7) # Optional: Add scatter points to show individual tag occurrences (same X for same participant) individual_tags = df[df["Tag_Word"] == target_tag] plt.scatter(x=individual_tags["Time_Order"], y=[0.3]*len(individual_tags), color="#ff7f0e", s=50, label="Individual Tag Use") # Customize labels and styling plt.xlabel("Participants (Chronological Order)", fontsize=12) plt.ylabel(f"Frequency of '{target_tag}'", fontsize=12) plt.title(f"'{target_tag}' Usage Across 10-Song Genre Annotation Participants", fontsize=14) plt.xticks(rotation=45) plt.legend() plt.tight_layout() plt.show()
If you want Y-axis to show the proportion of the target word relative to all tags a participant used, modify the frequency calculation like this:
# Calculate total tags per participant total_tags = df.groupby("Participant_ID")["Tagid"].count().reset_index(name="Total_Tags") # Calculate target word count per participant target_counts = df[df["Tag_Word"] == target_tag].groupby("Participant_ID")["Tagid"].count().reset_index(name="Target_Count") # Merge and compute proportion frequency_df = pd.merge(total_tags, target_counts, on="Participant_ID", how="left").fillna(0) frequency_df["Target_Proportion"] = frequency_df["Target_Count"] / frequency_df["Total_Tags"] # Merge with time order data frequency_df = frequency_df.merge(df[["Participant_ID", "Time_Order"]].drop_duplicates(), on="Participant_ID")
Then use Target_Proportion as your Y-axis value in the plot.
If you prefer R for your analysis, here's an equivalent setup:
library(ggplot2) library(dplyr) # Simulate data raw_data <- data.frame( Participant_ID = c("P1", "P1", "P2", "P3", "P3", "P3", "P4"), Tagid = c(1, 2, 1, 1, 2, 3, 1), Tag_Word = c("jazz", "blues", "jazz", "rock", "jazz", "jazz", "blues"), Time_Order = c(1, 1, 2, 3, 3, 3, 4) ) target_tag <- "jazz" # Calculate frequency frequency_df <- raw_data %>% filter(Tag_Word == target_tag) %>% group_by(Participant_ID) %>% summarise(Target_Count = n()) %>% left_join(raw_data %>% distinct(Participant_ID, Time_Order), by = "Participant_ID") # Plot ggplot(frequency_df, aes(x = factor(Time_Order), y = Target_Count)) + geom_bar(stat = "identity", fill = "#2ca02c", alpha = 0.7) + geom_jitter(data = raw_data %>% filter(Tag_Word == target_tag), aes(x = factor(Time_Order), y = 0.3), color = "#ff7f0e", size = 3, alpha = 0.7) + labs(x = "Participants (Chronological Order)", y = paste("Frequency of '", target_tag, "'", sep = ""), title = "'jazz' Usage Across 10-Song Genre Annotation Participants") + theme(axis.text.x = element_text(rotation = 45, hjust = 1))
内容的提问来源于stack exchange,提问作者jbouwman

