R语言:转换DataFrame字符变量大小写及保留str的技术问询
Hey there! Let's break down how to solve these two R data manipulation tasks clearly, with practical code examples you can use right away.
You have two solid approaches here—one using base R (no extra packages needed) and another using the tidyverse's dplyr for a more readable, pipe-based workflow.
Base R Method
First, identify which columns are character types, then apply toupper() or tolower() only to those columns:
# Sample DataFrame df <- data.frame( name = c("alice", "bob", "charlie"), age = c(25, 30, 35), city = c("london", "paris", "new york"), stringsAsFactors = FALSE ) # Find character columns char_cols <- sapply(df, is.character) # Convert to uppercase df[char_cols] <- lapply(df[char_cols], toupper) # Or convert to lowercase (just swap toupper() with tolower()) # df[char_cols] <- lapply(df[char_cols], tolower)
Tidyverse (dplyr) Method
If you prefer the tidyverse style, use mutate(across(...)) (recommended for dplyr 1.0.0+) to target character columns:
library(dplyr) # Convert to uppercase df_upper <- df %>% mutate(across(where(is.character), toupper)) # Convert to lowercase df_lower <- df %>% mutate(across(where(is.character), tolower))
Great news—both methods above work perfectly here, and they won’t alter your DataFrame’s structure (i.e., str() output stays identical). Since 6k rows is tiny for R, you won’t run into performance issues either.
The key is that we only modify character columns—all other column types (numeric, integer, factor, etc.) remain untouched, along with column order and metadata.
Example with Your Large DataFrame
Let’s say your DataFrame is named large_df:
# Base R approach char_cols_large <- sapply(large_df, is.character) large_df[char_cols_large] <- lapply(large_df[char_cols_large], toupper) # Dplyr approach (if you use tidyverse) large_df_upper <- large_df %>% mutate(across(where(is.character), toupper))
Verify the Structure Didn’t Change
Double-check that nothing shifted by comparing the original and modified DataFrame structures:
str(large_df) str(large_df_upper) # Should match exactly except for uppercase values in character columns
内容的提问来源于stack exchange,提问作者arunv

