R语言初学者求助:na.omit()删除NA值全量删行、自定义筛选无效的问题排查
Hey there! Let's figure out why your NA removal attempts aren't working as expected. I'll break down both problems and walk you through fixes:
1. Why na.omit() is deleting all rows
When na.omit() wipes out every single observation, it usually boils down to one of two scenarios:
- Every row has at least one standard NA: If even one column has NAs in every row, or every individual row has an NA in some column,
na.omit()will drop all rows—since it removes any row containing any NA value. - Your "missing values" aren't standard R NAs: Sometimes data uses placeholders like blank strings (
""), the character"NA", or other non-standard values thatis.na()doesn't recognize as true NAs.na.omit()only targets R's nativeNAtype, so it won't touch these.
How to diagnose and fix this:
First, check where the NAs (if any) are in your dataset:
# Total number of standard NAs across all columns sum(is.na(Year2021_Trips_v2)) # Count NAs per column to see which ones are problematic colSums(is.na(Year2021_Trips_v2))
If you see zero NAs here but know there are missing values, convert those non-standard placeholders to true NAs first:
# Turn blank strings into NAs Year2021_Trips_v2[Year2021_Trips_v2 == ""] <- NA # Turn character "NA" strings into NAs Year2021_Trips_v2[Year2021_Trips_v2 == "NA"] <- NA
Then you can either re-run na.omit(), or use a more targeted approach to only remove NAs from specific columns (instead of dropping rows with any NA):
# Using dplyr to keep rows without NAs in key columns Year2021_Trips_v2 <- Year2021_Trips_v2 %>% filter(!is.na(name_start_station), !is.na(ride_length)) # Add other columns as needed
2. Why your second code didn't remove any empty values
The code you wrote:
Year2021_Trips_v2 <- Year2021_Trips[!(Year2021_Trips$name_start_station == "HQ QR" | Year2021_Trips$ride_length < 0),]
doesn't include any logic to handle NA values at all! It only filters out rows where name_start_station is "HQ QR" or ride_length is negative. To add NA removal to this filter, you need to explicitly check for NAs:
# Base R version combining your rules with NA checks Year2021_Trips_v2 <- Year2021_Trips[!(Year2021_Trips$name_start_station == "HQ QR" | Year2021_Trips$ride_length < 0 | is.na(Year2021_Trips$name_start_station) | is.na(Year2021_Trips$ride_length)),] # Cleaner dplyr version Year2021_Trips_v2 <- Year2021_Trips %>% filter(name_start_station != "HQ QR", ride_length >= 0, !is.na(name_start_station), !is.na(ride_length))
This way, you're combining your original filtering rules with checks to exclude rows with NAs in the columns you care about.
内容的提问来源于stack exchange,提问作者Odiri Metieh

