R语言中gowdis函数计算样本距离返回NaN问题求助
gowdis() When Calculating Mixed-Type Data Distances It looks like you're hitting NaN values when using FD::gowdis() to compute distances between your samples—this is almost always tied to issues with data structure, missing values, or incompatible variable types. Let's break down the likely problems and fix them step by step.
1. Fix Mismatched Column Structures Between df1 and df2
Looking at your data example, df2 includes an ID column (the 1, 2, 3, 4 values) that df1 doesn't have. When you rbind() them, this creates missing values (NA) in the ID column for df1's row. gowdis() can't handle missing values, so it returns NaN for any distance involving that row.
Fix:
Either add an ID to df1 (and fill it appropriately) or remove the ID column from df2 to align the columns:
# Remove the ID column from df2 to match df1's structure df2 <- df2[, -1]
2. Ensure Correct Variable Types
gowdis() requires categorical variables (like tcp, http) to be factor types, and numerical columns to be numeric. If your categorical columns are stored as character strings, or numerical columns have non-numeric values, this can introduce hidden issues or NA values.
Fix:
Convert categorical columns to factors and verify numerical columns are numeric:
# Convert categorical columns to factors df1$proto <- as.factor(df1$proto) df1$service <- as.factor(df1$service) df2$proto <- as.factor(df2$proto) df2$service <- as.factor(df2$service) # Verify numerical columns are numeric (adjust column names to match your data) num_cols <- setdiff(colnames(df1), c("proto", "service")) df1[num_cols] <- lapply(df1[num_cols], as.numeric) df2[num_cols] <- lapply(df2[num_cols], as.numeric)
3. Check for Missing Values
Even after fixing column structure, hidden NA values can cause gowdis() to return NaN. Run this check on your merged data:
mixusefull2 <- rbind(df1, df2) anyNA(mixusefull2)
If this returns TRUE, locate and resolve the missing values (e.g., fill them with a reasonable default, or remove rows/columns if appropriate).
4. Revised Full Code
Here's the corrected workflow incorporating these fixes:
# Load required libraries library(FD) library(KernelKnn) # Fix column alignment (remove ID from df2) df2 <- df2[, -1] # Convert variables to correct types df1$proto <- as.factor(df1$proto) df1$service <- as.factor(df1$service) df2$proto <- as.factor(df2$proto) df2$service <- as.factor(df2$service) # Ensure numerical columns are numeric num_cols <- setdiff(colnames(df1), c("proto", "service")) df1[num_cols] <- lapply(df1[num_cols], as.numeric) df2[num_cols] <- lapply(df2[num_cols], as.numeric) # Merge data mixusefull2 <- rbind(df1, df2) # Check for missing values (resolve if any) if(anyNA(mixusefull2)) { warning("Missing values detected! Resolve before computing distances.") # Example fix: fill NA with mode for factors, median for numeric # mixusefull2 <- tidyr::fill(mixusefull2, everything()) } # Compute gowdis distance matrix dist_mixusefull2 <- as.matrix(gowdis(mixusefull2)) # Get kNN indices W <- nrow(df1) + 1 E <- nrow(df1) + nrow(df2) idxs_mixusefull2 <- distMat.knn.index.dist( dist_mixusefull2, TEST_indices = W:E, k = 1, threads = 1, minimize = TRUE )
Additional Notes
gowdis()uses Gower's distance, which handles mixed-type data, but it's sensitive to any missing values. Double-check your raw data for typos or incomplete entries.- If your numerical columns have very different scales, you can normalize them first (e.g., using
scale())—this won't fix NaNs, but it will improve the distance calculation's reliability.
内容的提问来源于stack exchange,提问作者maria

