如何删除数据框中含非法数据的行?附R语言代码尝试
Hey there! Let's get your code sorted out so you can filter out rows where the currency column has invalid entries (like numbers or currencies not in your approved list).
What's Wrong with Your Current Code?
Your line KS_2016.update$currency[KS_2016.update$currency =currency_list] has two key issues:
- You're using a single equals sign (
=), which is for assignment in R, not checking if a value belongs to a list. - This syntax only tries to modify the
currencycolumn, not filter out entire rows from the data frame.
Correct Solutions
We need to use the %in% operator to check if each value in currency is part of your approved list, then keep only those rows. Here are two common, reliable approaches:
1. Base R Approach
This uses core R syntax without extra packages:
# Define your valid currency list currency_list <- c("GBP","HKD","AUD","NZD","USD") # Filter the data frame to keep only rows with valid currencies KS_2016.update_cleaned <- KS_2016.update[KS_2016.update$currency %in% currency_list, ]
Don't miss the comma at the end of the index ([, ])—it ensures we keep all columns for the valid rows.
2. dplyr Approach (More Readable for Data Workflows)
If you use the dplyr package (a go-to for data cleaning), the code is even more intuitive:
library(dplyr) # Define valid currencies currency_list <- c("GBP","HKD","AUD","NZD","USD") # Filter rows with valid currencies KS_2016.update_cleaned <- KS_2016.update %>% filter(currency %in% currency_list)
This will automatically drop any rows where currency is a number, or any value not in your currency_list—exactly what you need!
Quick Verification
After running either code, you can double-check the result by listing unique values in the cleaned data frame:
unique(KS_2016.update_cleaned$currency)
You should only see the currencies from your approved list.
内容的提问来源于stack exchange,提问作者Hunter Knighton

