R语言table()函数返回重复因子问题求助——基于星战调查数据
table() Output Hey there! Let's break down why your table() result is showing repeated Anakin values like 0, 2, 0, 1, 5—even though you expect the values to be 1-5 plus NA. Here are the key reasons and fixes:
1. You aren't saving the factor conversion back to your data frame
When you run:
as.factor(starwars$Anakin) as.factor(starwars$Startrek)
This only creates a temporary factor object—it doesn't modify the original columns in your starwars data frame. Since you used stringsAsFactors = FALSE in read.csv2(), those columns stay as character vectors.
When table() processes a character vector, it treats every unique string as a separate category. If your Anakin column has strings that look identical but aren't (like "0" vs "0 " with a trailing space, or hidden whitespace carried over from Excel), table() will count them as distinct entries—hence the duplicate "0" or "2" rows you're seeing.
Fix this by assigning the factor conversion back to the columns:
starwars$Anakin <- as.factor(starwars$Anakin) starwars$Startrek <- as.factor(starwars$Startrek)
2. Empty strings from Excel aren't being treated as NA
You mentioned you set Excel's "N/A" to empty strings "", but in R, empty strings don't automatically become NA. So your starwars$Anakin column has "" instead of NA, which table() will display as a separate blank category. If you want to map those empty strings to NA (to match your expected "1-5 and NA" values), add this step before converting to factors:
# Convert empty strings to NA starwars$Anakin[starwars$Anakin == ""] <- NA
3. Possible data entry/import issues with the "0" values
You noted you expect Anakin values to be 1-5, but your table() shows 0s. This might stem from:
- Accidental 0 values in your Excel sheet (maybe some respondents were coded as 0 instead of 1)
- Import quirks: If Excel saved numbers as text with leading zeros, or if there was a formatting error during CSV export.
To check the exact unique values in your Anakin column (before converting to factors), run:
unique(starwars$Anakin)
This will reveal any hidden whitespace or unexpected entries that might be causing duplicates.
Once you apply these fixes, re-run your table() command, and you should see clean, non-duplicate factor levels matching your expected data.
内容的提问来源于stack exchange,提问作者Victor Galuppo

