如何用R语言移除特定字符串前数字并合并重复空格?
Hey there! Let's fix your character vector cleaning issues step by step.
Your original regex *. Votes didn't work because * is a quantifier in regex (matches the preceding element zero or more times), not a wildcard. That's why it ended up deleting everything—*. matches any character (.) zero or more times, paired with Votes it grabs the entire text up to the last occurrence of that string.
To target only numbers with thousand separators right before "Votes.", use this precise regex:
# Example text with the pattern text <- "123,456 Votes. STATE QUESTION NO. 1 Amendment to Title 15 of the Ne..." # Remove just the number (keep "Votes.") cleaned_num <- gsub("\\d+(?:,\\d+)* (?=Votes\\.)", "", text)
Breakdown of the regex:
\\d+: Matches one or more digits(?:,\\d+)*: Non-capturing group that matches zero or more instances of a comma followed by digits (handles thousand separators like123,456or1,234,567)(?=Votes\\.): Positive lookahead to ensure the number is immediately followed by "Votes." (avoids deleting other comma-containing numbers elsewhere)
If you actually want to remove the number and the " Votes." part (your original attempt hinted at this), adjust the regex to:
cleaned_num <- gsub("\\d+(?:,\\d+)* Votes\\.", "", text)
Once you've removed the target number, collapse any multiple consecutive spaces into a single one with:
final_text <- gsub("\\s+", " ", cleaned_num)
\\s+ matches one or more whitespace characters (spaces, tabs, etc.), replacing them with a single space.
Putting it all together:
# Sample input text <- "987,654 Votes. STATE QUESTION NO. 1 Amendment to Title 15 of the Ne..." # Step 1: Remove the number before "Votes." cleaned_num <- gsub("\\d+(?:,\\d+)* (?=Votes\\.)", "", text) # Step 2: Clean up spaces final_text <- gsub("\\s+", " ", cleaned_num) # Check the result cat(final_text) # Output: Votes. STATE QUESTION NO. 1 Amendment to Title 15 of the Ne...
内容的提问来源于stack exchange,提问作者ModalBro

