数据框中向量组合的频次统计及后续处理技术问询
Hey there! Let’s work through your three questions step by step, building on the progress you’ve already made with your dataframe and column combinations.
table() replacement issue The reason table() breaks your existing code is that it returns a contingency table object instead of a dataframe (which is what plyr::count() outputs). To fix this, you just need to convert the table result to a dataframe and align the column names to match what your workflow expects:
library(plyr) res_table <- list() for (j in 1:length(f)){ # Convert table output to dataframe tab_result <- as.data.frame(table(DF[, f[[j]]])) # Rename the frequency column to match count()'s "freq" (table uses "Freq" by default) colnames(tab_result)[ncol(tab_result)] <- "freq" res_table[[j]] <- tab_result } # Now res_table will have the same structure as your original res list res_table
This adjustment ensures the output matches the dataframe format your loop and downstream code rely on.
count() and table() To get a reliable performance comparison for your 10,000×8 dataframe, use the microbenchmark package—it runs each method multiple times to average out noise and gives you clear timing metrics. Here’s how to set it up:
# Install if you haven't already # install.packages("microbenchmark") library(microbenchmark) # Wrap your two methods into reusable functions count_based <- function() { res <- list() for (j in 1:length(f)){ res[[j]] <- count(DF, f[[j]]) } res } table_based <- function() { res <- list() for (j in 1:length(f)){ tab <- as.data.frame(table(DF[, f[[j]]])) colnames(tab)[ncol(tab)] <- "freq" res[[j]] <- tab } res } # Run the benchmark (adjust `times` based on how long you want it to run) benchmark_results <- microbenchmark( count_method = count_based(), table_method = table_based(), times = 100 ) # View the results print(benchmark_results) # Optional: Visualize the timing distribution library(ggplot2) autoplot(benchmark_results)
The output will show you metrics like median, mean, and min/max execution times, so you can see which method is faster for your specific dataset.
To organize your final results into frequency-based groups with sorted output, first combine all your results into a single dataframe (with a label for each column combo), then sort and group:
Step 1: Combine and label all results
# Add a column to track which column combo each row comes from labeled_res <- lapply(seq_along(res), function(i) { combo_label <- paste(f[[i]], collapse = "_") df <- res[[i]] df$combo <- combo_label df }) # Merge all list elements into one dataframe combined_results <- do.call(rbind, labeled_res)
Step 2: Sort and group by frequency
# Sort the combined dataframe by frequency (descending) and combo label sorted_results <- combined_results[order(-combined_results$freq, combined_results$combo), ] # Group rows by their frequency value grouped_by_freq <- split(sorted_results, sorted_results$freq) # Print formatted output (optional, for readability) for (freq in sort(names(grouped_by_freq), decreasing = TRUE)) { cat(sprintf("### Frequency = %s\n", freq)) print(grouped_by_freq[[freq]]) cat("\n") }
If you want to work only with results where freq > 1, just use your filtered list instead of the full res:
filtered_res <- lapply(res, subset, freq > 1) # Repeat the labeling/combining steps, skipping empty dataframes labeled_filtered <- lapply(seq_along(filtered_res), function(i) { if (nrow(filtered_res[[i]]) == 0) return(NULL) combo_label <- paste(f[[i]], collapse = "_") df <- filtered_res[[i]] df$combo <- combo_label df }) filtered_combined <- do.call(rbind, labeled_filtered) filtered_sorted <- filtered_combined[order(-filtered_combined$freq), ] filtered_grouped <- split(filtered_sorted, filtered_sorted$freq)
内容的提问来源于stack exchange,提问作者Citizen

