为数据框添加百分比列及非TOP10高频值频次的技术实现需求
Alright, let's work through your two requests: adding a percentage column to your top 10 frequency dataframe, and tallying up the total frequency of values that aren't in the top 10 range. Here's how to do it step by step based on your existing code:
1. Add Percentage Column to the Top 10 Dataframe
First, using the Z dataframe you already generated (with Value and Frequency columns), we can calculate what percentage each value's frequency makes up of the total top 10 frequency.
# Calculate total frequency of all top 10 values total_top10 <- sum(Z$Frequency) # Add a new "Percentage" column (rounded to 2 decimal places for readability) Z$Percentage <- round((Z$Frequency / total_top10) * 100, 2)
After running this, your Z dataframe will look something like this:
| Value | Frequency | Percentage |
|---|---|---|
| 1 | 635 | 41.23 |
| 0 | 296 | 19.32 |
| 1,000,000 | 115 | 7.51 |
| 10,000,000 | 110 | 7.18 |
| 20,000,000 | 104 | 6.80 |
| 5,000,000 | 101 | 6.60 |
| 50,000,000 | 86 | 5.61 |
| 25,000,000 | 85 | 5.54 |
| ... | ... | ... |
2. Calculate Total Frequency of Non-Top10 Values
To get the total number of entries that don't fall into your top 10 values, we can use the original frequency table and exclude the top 10 values:
# Get the full frequency table of your variable Y all_value_freq <- table(Y) # Extract the top 10 values from your Z dataframe (convert to character to match table names) top10_values <- as.character(Z$Value) # Sum the frequencies of all values NOT in the top 10 non_top10_total <- sum(all_value_freq[!names(all_value_freq) %in% top10_values]) # Optional: Add an "Other" row to your Z dataframe for complete summary other_row <- data.frame( Value = "Other", Frequency = non_top10_total, Percentage = round((non_top10_total / (total_top10 + non_top10_total)) * 100, 2) ) # Combine original top 10 with the "Other" row full_summary <- rbind(Z, other_row)
This full_summary dataframe will now include both your top 10 values and a consolidated row for all other values, with their respective frequencies and percentages relative to the entire dataset.
内容的提问来源于stack exchange,提问作者Bustergun

