R语言:为XLSX导入的首字母大写无空格列添加空格
Hey there! Let's sort out why your gsub call only works when you run it line-by-line, and get it working smoothly for your entire column from that XLSX file.
First, let's make sure we're importing your Excel data correctly. The readxl package is super reliable for XLSX files—here's how to use it:
# Install the package if you haven't already install.packages("readxl") library(readxl) # Import your file (replace "your_data.xlsx" with your actual file path) df <- read_excel("your_data.xlsx")
The Core Fix: Apply gsub to the Entire Column
Your regex pattern is actually spot-on! The issue is likely that you weren't applying it to the full column vector, or your column was stored as a factor instead of plain text. Here's how to fix both cases:
Step 1: Ensure your column is character type
If your target column (let's say it's named camel_col) is a factor, gsub will act on the factor levels instead of the actual text. Convert it to character first:
df$camel_col <- as.character(df$camel_col)
Step 2: Run gsub on the entire column
Now you can apply your regex to every row in one go:
# Create a new column with space-separated text df$space_separated <- gsub("([[:upper:]])([[:upper:]][[:lower:]])", "\\1 \\2", df$camel_col)
Let's test this with your examples:
- Input:
HowDoYouWorkOnThis→ Output:How Do You Work On This - Input:
ThisIsGreatExample→ Output:This Is Great Example
Perfect, that's exactly what you want!
Bonus: Tidyverse Syntax (If You Prefer)
If you use the tidyverse, dplyr makes this even cleaner:
install.packages("dplyr") library(dplyr) df <- df %>% mutate( camel_col = as.character(camel_col), space_separated = gsub("([[:upper:]])([[:upper:]][[:lower:]])", "\\1 \\2", camel_col) )
Why It Only Worked Line-by-Line Before?
Chances are one of these was the culprit:
- Your column was a factor instead of character (so
gsubcouldn't process the text directly) - You were manually applying
gsubto individual rows (likedf$camel_col[1]) instead of the full column vector
Quick Test to Confirm
If you want to verify the regex works on a sample vector first:
test_cases <- c("HowDoYouWorkOnThis", "ThisIsGreatExample") gsub("([[:upper:]])([[:upper:]][[:lower:]])", "\\1 \\2", test_cases)
This will instantly return the space-separated versions without any line-by-line work.
That should solve your problem—now you can convert the entire column in one step instead of processing each row manually!
内容的提问来源于stack exchange,提问作者Praveen M Kulkarni

