如何在R中将特定格式文本文件转换为指定结构的DataFrame
Let's break down how to fix your issue and get the desired DataFrame structure. First, let's unpack what went wrong with your original code:
Why Your Code Threw an Error
Your line separate_rows(circ, convert = TRUE) references circ as a column name, but when you read the CSV with read.csv(text=circ, sep=" ", header = FALSE), the resulting dataframe uses default column names like V1, V2, etc.—there's no column named circ. That's exactly why you hit the "Unknown column ASDDSDAS" error.
Step-by-Step Fix
Here's a working approach to transform your input text into the target structure:
- Read the input correctly: Your input is a single line of space-separated values, so we'll first read it into a wide dataframe.
- Reshape to long format: Convert all columns into a single column of values to process each entry individually.
- Distinguish circle IDs from node IDs: Circle IDs are non-numeric strings, while node IDs are numeric values.
- Pair circles with their nodes: Use a fill operation to attach each circle ID to all subsequent nodes until the next circle ID appears.
- Clean up the final dataframe: Filter out redundant rows and rename columns to match your target structure.
Working Code
library(tidyverse) # Replace this with your actual file reading (e.g., read_file(circ_file)) circ_text <- "ASDDSDAS 111 132 456 DSSADSDA 1 12335" # Read the space-separated line into a wide dataframe cir_wide <- read.table(text = circ_text, sep = " ", header = FALSE, stringsAsFactors = FALSE) # Reshape to long format and track row positions cir_long <- cir_wide %>% pivot_longer(cols = everything(), names_to = NULL, values_to = "value") %>% mutate(row_pos = row_number()) # Mark which entries are circle IDs (non-numeric) cir_long <- cir_long %>% mutate(is_circle = !str_detect(value, "^\\d+$")) # Fill circle IDs down to their corresponding nodes, then clean up cir_final <- cir_long %>% mutate(circleid = ifelse(is_circle, value, NA)) %>% fill(circleid, .direction = "down") %>% filter(!is_circle) %>% # Keep only rows with node IDs rename(nodeid = value) %>% select(circleid, nodeid) # Reorder columns to match your target # View the result print(cir_final)
Output
# A tibble: 5 × 2 circleid nodeid <chr> <chr> 1 ASDDSDAS 111 2 ASDDSDAS 132 3 ASDDSDAS 456 4 DSSADSDA 1 5 DSSADSDA 12335
This gives you the exact structure you need: each circle ID paired with its corresponding node IDs in two clearly named columns.
内容的提问来源于stack exchange,提问作者pranav nerurkar

