为何Python修改Spotfire列属性比R快?求R代码优化方案
Optimizing R Code for Spotfire Column Property Modifications
Hey there, let's dig into why your R code is running so much slower than Python when updating Spotfire column properties—24 seconds vs 4 seconds is a big gap, and we can fix that. I’ve looked over your code snippets, and there are several key bottlenecks we can address to get R’s runtime closer to Python’s speed.
Key Bottlenecks in Your Original R Code
- Repeated Attribute Lookups: Every time you call
attributes(temp2[,col])$SpotfireColumnMetaData, R has to re-traverse the data frame’s attribute hierarchy from scratch. Doing this multiple times per column adds up quickly. - Unnecessary Per-Column Processing: The
lapply(names(...), tolower)in your else block runs once for every column after the 4th—this only needs to be done once for all relevant columns indt. - Global Variable Overhead: Your
lapplyversion uses<<-to modify a globalcountvariable, which is slow and introduces unnecessary side effects. - Inefficient Column Indexing: Using
temp2[,col]in loops forces R to extract the entire column each time, even though you only need to modify its attributes.
Optimized R Code
We’ll pre-process reusable components and minimize repeated attribute lookups to cut down on runtime:
# Start timing start <- Sys.time() # Pre-cache Spotfire metadata for all columns in temp2 (avoids repeated lookups) temp2_meta <- lapply(temp2, function(col) attributes(col)$SpotfireColumnMetaData) # Pre-process dt's metadata: convert names to lowercase ONCE for all relevant columns dt_meta <- lapply(dt[, 1:(ncol(temp2)-4)], function(col) { meta <- attributes(col)$SpotfireColumnMetaData names(meta) <- tolower(names(meta)) meta }) # Handle first 4 columns for (col in 1:4) { meta <- temp2_meta[[col]] meta$upper <- Inf meta$lower <- -Inf meta$upper2 <- Inf meta$lower2 <- -Inf # Assign modified metadata back to the column attributes(temp2[[col]])$SpotfireColumnMetaData <- meta } # Handle remaining columns (starting from 5th) for (col in 5:ncol(temp2)) { dt_col_idx <- col - 4 meta <- temp2_meta[[col]] dt_meta_col <- dt_meta[[dt_col_idx]] meta$upper <- 2 meta$lower <- 1 meta$upper2 <- dt_meta_col$upper meta$lower2 <- dt_meta_col$lower # Assign modified metadata back attributes(temp2[[col]])$SpotfireColumnMetaData <- meta } # Print runtime cat("Runtime:", difftime(Sys.time(), start, units = "secs"), "seconds\n")
Even Faster: Vectorized Batch Updates
To push performance further, we can avoid per-column loops for the first 4 columns by updating them in a batch:
start <- Sys.time() # Pre-cache metadata lists temp2_meta <- lapply(temp2, function(col) attributes(col)$SpotfireColumnMetaData) dt_meta <- lapply(dt[, 1:(ncol(temp2)-4)], function(col) { meta <- attributes(col)$SpotfireColumnMetaData names(meta) <- tolower(names(meta)) meta }) # Batch update first 4 columns temp2_meta[1:4] <- lapply(temp2_meta[1:4], function(meta) { meta$upper <- Inf meta$lower <- -Inf meta$upper2 <- Inf meta$lower2 <- -Inf meta }) # Update remaining columns with mapply (vectorized loop alternative) temp2_meta[5:ncol(temp2)] <- mapply(function(meta, dt_meta_col) { meta$upper <- 2 meta$lower <- 1 meta$upper2 <- dt_meta_col$upper meta$lower2 <- dt_meta_col$lower meta }, temp2_meta[5:ncol(temp2)], dt_meta, SIMPLIFY = FALSE) # Assign all modified metadata back to temp2 in one pass for (i in seq_along(temp2_meta)) { attributes(temp2[[i]])$SpotfireColumnMetaData <- temp2_meta[[i]] } cat("Runtime:", difftime(Sys.time(), start, units = "secs"), "seconds\n")
Why This Works
- Pre-Caching: We pull all metadata objects once at the start, eliminating repeated traversal of the attribute hierarchy.
- Single Name Conversion: Converting
dt's metadata names to lowercase happens once, not per column. - Efficient Column Access: Using
temp2[[col]]instead oftemp2[,col]is faster for accessing individual columns and their attributes. - Batch Processing: The vectorized approach reduces loop overhead by handling multiple columns at once.
内容的提问来源于stack exchange,提问作者Nikita Belooussov
相关产品推荐
相关产品推荐

