R语言读取UCI成人数据集DataFrame的收入列问题咨询
Fixing & Troubleshooting Your Adult Dataset Read in R
Hey there! Let's start by cleaning up your code to avoid common reading issues, then walk through the most frequent problems folks run into with this dataset.
First, Here's a Polished Version of Your Loading Code
# Define the dataset URL url <- "http://archive.ics.uci.edu/ml/machine-learning-databases/adult/adult.data" # Read the CSV with proper handling adult <- read.csv( url, strip.white = TRUE, header = FALSE, na.strings = "?", # Critical: the dataset uses "?" for missing values stringsAsFactors = FALSE # Optional: keep strings as character vectors (default in R 4.0+) ) # Clean up column names (remove extra spaces, use snake_case for easier referencing) colnames(adult) <- c( "age", "workclass", "final_weight", "education", "education_num", "marital_status", "occupation", "relationship", "race", "sex", "capital_gain", "capital_loss", "hours_per_week", "native_country", "income" )
Common Reading Issues & Fixes
- Extra spaces in column names: Your original code had spaces around names like
" workclass "— this makes referencing columns (e.g.,adult$workclass) throw errors. The polished code removes those spaces and uses snake_case, which is standard for R. - Unrecognized missing values: The Adult dataset uses
?to mark missing entries. Withoutna.strings = "?", these will be treated as regular character values instead ofNA, which breaks downstream analysis likesummary()or modeling. - Unexpected factor columns: If you're using an older R version (pre-4.0),
read.csv()defaults to converting strings to factors. AddingstringsAsFactors = FALSEkeeps them as character vectors, which is usually more flexible unless you explicitly want factors. - Leading/trailing whitespace in data: You already used
strip.white = TRUE— great call! This removes extra spaces from individual data points (like in theoccupationornative_countrycolumns).
Quick Check to Verify Loading Worked
Run these commands to confirm everything looks right:
# View first 5 rows head(adult) # Check for missing values colSums(is.na(adult)) # Inspect the income column (your target variable) table(adult$income)
内容的提问来源于stack exchange,提问作者Lucas Fuzato
相关产品推荐
相关产品推荐

