如何基于3个情感因子水平构建0-100区间的负面累积指数
Let's walk through this step by step—this approach will give you a clear, meaningful index that aligns exactly with your requirements:
Step 1: Clean Sentiment Data & Map to Scores
First, we need to handle that empty factor level in your data and map each valid sentiment to its corresponding score (3 for negative, 2 for neutral, 1 for positive). We'll turn empty strings into NA since they aren't part of your defined sentiment levels.
library(dplyr) # Load your data (assuming it's already in a data frame called `data`) # Map sentiment levels to scores, handling empty strings as NA score_mapping <- c("negative" = 3, "neutral" = 2, "positive" = 1) data$sentiment_score <- score_mapping[data$sentiment] # Optional: Filter out rows with NA sentiment if you don't want to include them # data <- data %>% filter(!is.na(sentiment_score))
Step 2: Calculate Daily Cumulative Scores
Next, we'll group your data by day and compute two key values for each day:
- The total sum of sentiment scores (your "cumulative result" for the day)
- The number of valid sentiment observations (to use for normalization)
daily_summary <- data %>% group_by(date) %>% # Replace `date` with your actual day column name summarise( total_daily_score = sum(sentiment_score, na.rm = TRUE), valid_obs = sum(!is.na(sentiment_score)) # Count of non-NA sentiment entries )
Step 3: Normalize to 0-100 Range
This is the critical part—we'll normalize each day's score relative to its theoretical minimum and maximum possible values. This ensures 100 means all observations that day are negative, 0 means all are positive, and neutral sits at 50. We'll also handle days with no valid observations to avoid division by zero.
daily_summary <- daily_summary %>% mutate( # Theoretical min (all positive) and max (all negative) scores for the day min_possible = valid_obs * 1, max_possible = valid_obs * 3, # Calculate normalized index normalized_negative_index = case_when( valid_obs == 0 ~ NA_real_, # Mark days with no data as NA TRUE ~ ((total_daily_score - min_possible) / (max_possible - min_possible)) * 100 ) )
Why This Works:
- Per-day normalization: Unlike global normalization, this accounts for varying numbers of observations per day. A day with 5 all-negative entries will hit 100, just like a day with 100 all-negative entries—fairly comparing sentiment intensity across days.
- Clear interpretation: The index directly reflects how far the day's sentiment leans toward negative (100) or positive (0). Neutral days will land at 50, which makes intuitive sense.
Example Output
If you run head(daily_summary), you'll see something like this:
| date | total_daily_score | valid_obs | min_possible | max_possible | normalized_negative_index |
|---|---|---|---|---|---|
| 2024-01-01 | 15 | 5 | 5 | 15 | 100.0 |
| 2024-01-02 | 7 | 5 | 5 | 15 | 20.0 |
| 2024-01-03 | 10 | 5 | 5 | 15 | 50.0 |
In this example:
- Jan 1 has all negative sentiment (100)
- Jan 2 has 3 positive + 2 neutral (20, low negative)
- Jan3 has all neutral (50, midpoint)
Optional: Global Normalization (If Needed)
If you prefer to normalize across all days (instead of per day), use this code instead of Step3. Note this is less ideal for comparing intensity, but it's an option:
# Get global min and max scores across all days global_min <- min(daily_summary$total_daily_score, na.rm = TRUE) global_max <- max(daily_summary$total_daily_score, na.rm = TRUE) daily_summary <- daily_summary %>% mutate( global_normalized_index = ((total_daily_score - global_min) / (global_max - global_min)) * 100 )
内容的提问来源于stack exchange,提问作者Kash

