认知地图累积曲线绘制技术问询:基于40份认知地图的新增变量累积统计
Got it, let's walk through exactly how to build this accumulation curve for your cognitive maps. I'll cover both Python and R solutions since those are the most common tools for this kind of data analysis.
Python Solution (Using Pandas & Matplotlib)
Step 1: Track New Variables Across Maps
First, we'll use a set to keep track of variables we've already seen, then iterate through each map to count how many new variables it contributes:
import pandas as pd import matplotlib.pyplot as plt # Assume your dataframe is named `df` with row indices 1-40 and columns F1-F144 seen_variables = set() new_vars_per_map = [] # Iterate over each map (row) for map_num, row in df.iterrows(): # Get all variables present in the current map (where value is 1) current_vars = row[row == 1].index.tolist() # Filter variables that haven't been seen before new_vars = [var for var in current_vars if var not in seen_variables] # Record the count of new variables new_vars_per_map.append(len(new_vars)) # Update our set of seen variables seen_variables.update(new_vars) # Convert results to a Series for easier plotting new_vars_series = pd.Series(new_vars_per_map, index=df.index, name="New Variables")
Step 2: Plot the Accumulation Curve
Now we can plot both the per-map new variable count (your requested Y-axis) and the cumulative total of unique variables (a common complementary plot):
# Plot the per-map new variables curve plt.figure(figsize=(10, 6)) plt.plot(new_vars_series.index, new_vars_series.values, marker="o", linestyle="-", color="#1f77b4") plt.xlabel("Map Number") plt.ylabel("Number of New Variables") plt.title("Accumulation Curve: New Variables per Cognitive Map") plt.grid(True, alpha=0.3) plt.show() # Optional: Plot cumulative total of unique variables cumulative_total = new_vars_series.cumsum() plt.figure(figsize=(10, 6)) plt.plot(cumulative_total.index, cumulative_total.values, marker="s", linestyle="--", color="#ff7f0e") plt.xlabel("Map Number") plt.ylabel("Total Unique Variables") plt.title("Cumulative Total of Unique Variables Across Maps") plt.grid(True, alpha=0.3) plt.show()
R Solution (Using Base R & ggplot2)
Step 1: Calculate New Variables per Map
We'll use a vector to track seen variables and loop through each row to count new entries:
# Assume your dataframe is named `df` with row names 1-40 and columns F1-F144 seen_variables <- c() new_vars_per_map <- numeric(nrow(df)) # Loop through each map for (i in 1:nrow(df)) { # Get variables present in the current map current_vars <- colnames(df)[df[i, ] == 1] # Find variables not seen before new_vars <- setdiff(current_vars, seen_variables) # Record the count new_vars_per_map[i] <- length(new_vars) # Update seen variables seen_variables <- union(seen_variables, new_vars) } # Create a dataframe for plotting plot_data <- data.frame( Map_Number = as.integer(rownames(df)), New_Variables = new_vars_per_map )
Step 2: Generate the Plot
Use either base R or ggplot2 for visualization:
# Base R plot plot(plot_data$Map_Number, plot_data$New_Variables, type = "o", col = "#1f77b4", xlab = "Map Number", ylab = "Number of New Variables", main = "Accumulation Curve: New Variables per Cognitive Map", grid = TRUE) # ggplot2 plot (more customizable) library(ggplot2) ggplot(plot_data, aes(x = Map_Number, y = New_Variables)) + geom_line(color = "#1f77b4") + geom_point(shape = 16, color = "#1f77b4") + labs(x = "Map Number", y = "Number of New Variables", title = "Accumulation Curve: New Variables per Cognitive Map") + theme_minimal() + theme(panel.grid.major = element_line(alpha = 0.3)) # Optional: Cumulative total plot plot_data$Cumulative_Total <- cumsum(plot_data$New_Variables) ggplot(plot_data, aes(x = Map_Number, y = Cumulative_Total)) + geom_line(color = "#ff7f0e", linetype = "dashed") + geom_point(shape = 17, color = "#ff7f0e") + labs(x = "Map Number", y = "Total Unique Variables", title = "Cumulative Total of Unique Variables Across Maps") + theme_minimal()
Quick Notes
- Data Checks: Make sure your row indices/names are actually integers 1-40. If not, convert them first (e.g.,
df.index = df.index.astype(int)in Python,rownames(df) <- as.integer(rownames(df))in R). - Missing Values: If you have
NAvalues in your dataframe, replace them with 0 first (e.g.,df.fillna(0, inplace=True)in Python,df[is.na(df)] <- 0in R). - Saturation Insight: The cumulative total plot is great for seeing when you stop discovering new variables (the curve flattens out), which can indicate you've captured most of the relevant variables in your maps.
内容的提问来源于stack exchange,提问作者Gwyn
相关产品推荐
相关产品推荐

