You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

数据框中向量组合的频次统计及后续处理技术问询

Hey there! Let’s work through your three questions step by step, building on the progress you’ve already made with your dataframe and column combinations.

1. Fixing the table() replacement issue

The reason table() breaks your existing code is that it returns a contingency table object instead of a dataframe (which is what plyr::count() outputs). To fix this, you just need to convert the table result to a dataframe and align the column names to match what your workflow expects:

library(plyr)
res_table <- list()
for (j in 1:length(f)){
  # Convert table output to dataframe
  tab_result <- as.data.frame(table(DF[, f[[j]]]))
  # Rename the frequency column to match count()'s "freq" (table uses "Freq" by default)
  colnames(tab_result)[ncol(tab_result)] <- "freq"
  res_table[[j]] <- tab_result
}

# Now res_table will have the same structure as your original res list
res_table

This adjustment ensures the output matches the dataframe format your loop and downstream code rely on.

2. Comparing execution time between count() and table()

To get a reliable performance comparison for your 10,000×8 dataframe, use the microbenchmark package—it runs each method multiple times to average out noise and gives you clear timing metrics. Here’s how to set it up:

# Install if you haven't already
# install.packages("microbenchmark")
library(microbenchmark)

# Wrap your two methods into reusable functions
count_based <- function() {
  res <- list()
  for (j in 1:length(f)){
    res[[j]] <- count(DF, f[[j]])
  }
  res
}

table_based <- function() {
  res <- list()
  for (j in 1:length(f)){
    tab <- as.data.frame(table(DF[, f[[j]]]))
    colnames(tab)[ncol(tab)] <- "freq"
    res[[j]] <- tab
  }
  res
}

# Run the benchmark (adjust `times` based on how long you want it to run)
benchmark_results <- microbenchmark(
  count_method = count_based(),
  table_method = table_based(),
  times = 100
)

# View the results
print(benchmark_results)

# Optional: Visualize the timing distribution
library(ggplot2)
autoplot(benchmark_results)

The output will show you metrics like median, mean, and min/max execution times, so you can see which method is faster for your specific dataset.

3. Grouping and sorting results by frequency

To organize your final results into frequency-based groups with sorted output, first combine all your results into a single dataframe (with a label for each column combo), then sort and group:

Step 1: Combine and label all results

# Add a column to track which column combo each row comes from
labeled_res <- lapply(seq_along(res), function(i) {
  combo_label <- paste(f[[i]], collapse = "_")
  df <- res[[i]]
  df$combo <- combo_label
  df
})

# Merge all list elements into one dataframe
combined_results <- do.call(rbind, labeled_res)

Step 2: Sort and group by frequency

# Sort the combined dataframe by frequency (descending) and combo label
sorted_results <- combined_results[order(-combined_results$freq, combined_results$combo), ]

# Group rows by their frequency value
grouped_by_freq <- split(sorted_results, sorted_results$freq)

# Print formatted output (optional, for readability)
for (freq in sort(names(grouped_by_freq), decreasing = TRUE)) {
  cat(sprintf("### Frequency = %s\n", freq))
  print(grouped_by_freq[[freq]])
  cat("\n")
}

If you want to work only with results where freq > 1, just use your filtered list instead of the full res:

filtered_res <- lapply(res, subset, freq > 1)

# Repeat the labeling/combining steps, skipping empty dataframes
labeled_filtered <- lapply(seq_along(filtered_res), function(i) {
  if (nrow(filtered_res[[i]]) == 0) return(NULL)
  combo_label <- paste(f[[i]], collapse = "_")
  df <- filtered_res[[i]]
  df$combo <- combo_label
  df
})

filtered_combined <- do.call(rbind, labeled_filtered)
filtered_sorted <- filtered_combined[order(-filtered_combined$freq), ]
filtered_grouped <- split(filtered_sorted, filtered_sorted$freq)

内容的提问来源于stack exchange,提问作者Citizen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 04:06:05