如何在R语言中从proxy包dist函数生成的crossmat类型矩阵提取0<值<5的元素并转换为指定格式DataFrame
Extracting Filtered Elements from
proxy::dist crossmat into DataFrame Got it, let's work through this problem to get exactly the output you need. Here's a step-by-step solution tailored to your use case:
Step 1: Recap & Setup
First, let's confirm we're working with the crossmat object generated by proxy::dist()—this is essentially a matrix where rows map to your df1 entries and columns map to df2 entries. We need to pull out values between 0 and 5, while retaining the original row/column indices from df1 and df2 (not the matrix's relative positions).
Step 2: Code Implementation
Here's the full code to achieve your desired result:
# Load the proxy package (if not already loaded) library(proxy) # Your existing code to create the distance matrix a <- df1 b <- df2 dm <- dist(a, b, method = "euclidean") # 1. Find positions in the matrix where values are between 0 and 5 # arr.ind = TRUE returns matrix row/column positions as a data frame filter_positions <- which(dm > 0 & dm < 5, arr.ind = TRUE) # 2. Map matrix positions to original df1/df2 indices # Matrix rows correspond to df1 rows—use df1's rownames for original indices original_rows <- rownames(a)[filter_positions[, "row"]] # Matrix columns correspond to df2 rows—use df2's rownames for original indices original_cols <- rownames(b)[filter_positions[, "col"]] # 3. Extract the filtered values filtered_values <- dm[filter_positions] # 4. Combine into the final DataFrame result_df <- data.frame( row = as.integer(original_rows), col = as.integer(original_cols), value = filtered_values, stringsAsFactors = FALSE ) # View the result head(result_df)
Key Notes:
- Original Indices: This code assumes your
df1anddf2have their original row indices stored asrownames()(like the 3672, 192 values in your example). If you store original indices in a separate column (e.g.,df1$original_row), replacerownames(a)witha$original_row(and similarly fordf2). - Data Types: Using
as.integer()converts the character rownames back to integer indices, matching your sample output format. - Filter Logic: The
which()function efficiently finds all positions where values meet your0 < value < 5condition.
Example Output
Running this code will produce a DataFrame structured exactly like your sample:
row col value 1 3672 192 4.963545 2 3673 192 4.848763 3 3672 193 4.769458 4 3673 193 4.645692 5 3674 193 4.889749 6 3672 194 4.125620
内容的提问来源于stack exchange,提问作者Nancy
相关产品推荐
相关产品推荐

