如何导入变量定义在外部文件的固定宽度文件到R?YRBS数据导入求助
Hey there! Let's walk through how to import that YRBS fixed-width data into R, especially since we need to prep it for the survey package. This is a super common task with CDC datasets, so let's break it down step by step:
First, make sure you have two key files saved locally:
- The fixed-width data file (typically a
.dator.txtfile from the YRBS page) - The codebook/metadata file (this tells you variable names, start/end positions, value labels, and sampling details like weights, strata, and PSUs)
Fixed-width files depend on knowing exactly where each variable starts and ends. Instead of manually transcribing this, we can extract the specs from the codebook:
If your codebook is a text file:
Use base R to load and parse it with regex (adjust the regex to match your codebook's formatting—YRBS codebooks usually have consistent lines with variable name, start, end):
# Load the codebook text codebook_lines <- readLines("yrbs_codebook.txt") # Extract lines with variable position info var_spec_lines <- codebook_lines[grepl("^[A-Z0-9]+\\s+[0-9]+\\s+[0-9]+", codebook_lines)] # Split each line into parts var_specs <- strsplit(var_spec_lines, "\\s+") # Pull out names, start positions, end positions var_names <- sapply(var_specs, `[`, 1) var_starts <- as.integer(sapply(var_specs, `[`, 2)) var_ends <- as.integer(sapply(var_specs, `[`, 3))
If your codebook is a PDF:
Use the pdftools package to convert it to text first:
library(pdftools) # Convert PDF to text codebook_text <- pdf_text("yrbs_codebook.pdf") codebook_lines <- unlist(strsplit(codebook_text, "\n")) # Then follow the same regex extraction steps as above
With the variable specs ready, use read.fwf() to load the data correctly:
# Calculate column widths (end position - start position + 1) var_widths <- var_ends - var_starts + 1 # Import the data yrbs_raw <- read.fwf( file = "yrbs_data.dat", widths = var_widths, col.names = var_names, stringsAsFactors = FALSE, encoding = "ASCII" # YRBS files often use ASCII encoding )
Since YRBS uses complex sampling, we need to create a survey design object with the survey package. First, confirm the weight, stratum, and PSU variable names from the codebook (they're usually something like WEIGHT, STRATUM, PSU):
library(survey) # Create the survey design yrbs_survey <- svydesign( id = ~PSU, strata = ~STRATUM, weights = ~WEIGHT, data = yrbs_raw, nest = TRUE # YRBS uses nested sampling, so this is important )
YRBS variables use coded values (e.g., 1 = Yes, 2 = No). The labelled package makes it easy to add these labels for clarity:
library(labelled) # Example: Add labels to a smoking variable yrbs_raw$smoke <- set_value_labels( yrbs_raw$smoke, `1` = "Yes", `2` = "No", `9` = "Don't know" )
- Column mismatch errors: Double-check your start/end positions—sometimes codebooks use 1-indexing, which is what
read.fwf()expects, but if you're getting issues, confirm the positions again. - Garbled text: If you see weird characters, make sure you're using the correct encoding (ASCII works for most YRBS files).
- Survey design errors: Ensure weight, stratum, and PSU variables are numeric. If they're stored as characters, convert them with
as.integer()oras.numeric().
内容的提问来源于stack exchange,提问作者Chris

