R语言函数开发:生成指定排序的命名数值列表(基于查找表)
Hey there! Let's build this R function step by step to meet all your requirements. First, let's clarify the core needs, then dive into the code and adjustments.
需求回顾
- The output
q.arg2must match the length of the input character vectorfactors2 - For each name in
factors2:- If its distribution type in
dist2is"normal", extract the correspondingmeanandsdvalues fromdist2and return a sub-list likelist(mean=x, sd=y) - If it's another distribution (e.g.,
"weibull"), calculate theshapeandscaleparameters using the corresponding observations in data framek2, then returnlist(shape=x, scale=y)
- If its distribution type in
- Finally, sort
q.arg2according to the order specified insort2
修正参考数据结构
First, let's fix the dist2 data frame from your reference code—its original row-based structure isn't easy to query. We'll restructure it so each variable (var1/var2/var3) is a row with its corresponding id, mean, sd, and distribution:
# Generate dist2 correctly id <- c("a", "b", "c") mean <- c(3, 1, 4) sd <- c(2,7,8) distribution <- c("weibull", "normal", "normal") dist2 <- data.frame(id, mean, sd, distribution) rownames(dist2) <- c("var1", "var2", "var3") # Map rows to variable names dist2$mean <- as.numeric(dist2$mean) dist2$sd <- as.numeric(dist2$sd) # Generate factors2 factors2 <- c("var2", "var3") # Generate k2 (data for weibull fitting, fixed rnorm parameter order) set.seed(123) # For reproducible results k2 <- data.frame(replicate(2, rnorm(n=300, mean=100, sd=2))) colnames(k2) <- c("var1", "var2") # Load required package library(fitdistrplus) # Sort order (matches your reference "a,b" which maps to var1/var2) sort2 <- "a,b"
完整函数实现
This function handles both variable-name and id-based sorting, and includes error handling for missing variables or unsupported distributions:
generate_qarg2 <- function(factors2, dist2, k2, sort2) { # Handle sort2: if it uses ids (like "a,b"), map to variable names sort_ids <- strsplit(sort2, ",")[[1]] sort_order <- rownames(dist2)[match(sort_ids, dist2$id)] # Initialize empty list to store results q.arg2 <- list() # Iterate over each variable in factors2 for (var in factors2) { # Get distribution info for the current variable dist_info <- dist2[var, ] if (dist_info$distribution == "normal") { # Extract pre-defined mean and sd (ensured numeric) q.arg2[[var]] <- list(mean = dist_info$mean, sd = dist_info$sd) } else if (dist_info$distribution == "weibull") { # Check if variable exists in k2 before fitting if (!var %in% colnames(k2)) { stop(paste("Variable", var, "not found in k2 data frame")) } # Fit weibull distribution and extract parameters fw <- fitdist(k2[[var]], "weibull") q.arg2[[var]] <- list(shape = fw$estimate["shape"], scale = fw$estimate["scale"]) } else { # Handle unsupported distribution types stop(paste("Unsupported distribution type:", dist_info$distribution)) } } # Sort the list: keep variables in sort_order first, then add remaining ones sorted_vars <- intersect(sort_order, factors2) sorted_vars <- c(sorted_vars, setdiff(factors2, sorted_vars)) q.arg2 <- q.arg2[sorted_vars] # Remove top-level names to match your reference q.arg2 format names(q.arg2) <- NULL return(q.arg2) }
函数使用示例
# Run the function with your sample data q.arg2_result <- generate_qarg2(factors2, dist2, k2, sort2) # View the output print(q.arg2_result)
Sample output (matches your reference format):
[[1]] [[1]]$mean [1] 1 [[1]]$sd [1] 7 [[2]] [[2]]$mean [1] 4 [[2]]$sd [1] 8
代码解释
- Data Structure Fix: We restructured
dist2to make variable-specific queries straightforward, and convertedmean/sdto numeric types (they were stored as characters in your original code). - Distribution Handling: For normal distributions, we pull pre-defined parameters directly from
dist2. For weibull distributions, we usefitdist()fromfitdistrplusto estimate shape/scale fromk2. - Sorting Logic: The function maps
sort2ids to variable names (matching your reference"a,b"input), then reorders the result list to match this order. Any variables infactors2not insort2are added to the end. - Error Handling: Includes checks for missing variables in
k2and unsupported distribution types to avoid unexpected crashes.
内容的提问来源于stack exchange,提问作者Avery
相关产品推荐
相关产品推荐

