如何用R正则修正网页字符串的标点与缩写格式?
Hey there! Let's sort out that annoying formatting problem with your parsed string. The core issue is we need to eliminate unnecessary spaces right before punctuation marks—whether it's apostrophes in contractions (like 's or hasn't) or commas/periods—without accidentally deleting the punctuation itself.
The Problem with Your Original Approach
Your initial regex str_remove_all(x,"\\s[[:punct:]]'") was too narrow and risked removing punctuation because it targeted specific combinations (space + punctuation + apostrophe) instead of addressing the root cause: any space preceding a punctuation mark.
The Solution: Target Spaces Before Punctuation
Use str_replace_all() (from the stringr package) with a regex that captures the punctuation mark after the space, then replaces the entire space-punctuation pair with just the punctuation. Here's how:
Original String:
x = "the company 's newly launched cryptocurrency , Libra , hasn 't been contacted by Facebook , according to a report ."Fixing Code:
library(stringr) fixed_string <- str_replace_all(x, "\\s+([[:punct:]])", "\\1")Result:
"the company's newly launched cryptocurrency, Libra, hasn't been contacted by Facebook, according to a report."
How This Works
\\s+: Matches one or more whitespace characters (handles single or multiple accidental spaces)([[:punct:]]): Captures any punctuation mark (apostrophe, comma, period, etc.) into a group\\1: Replaces the matched space-punctuation pair with just the captured punctuation mark
This regex is flexible enough to fix all your spacing issues in one go, whether it's contractions like hasn 't → hasn't or misplaced spaces around commas like cryptocurrency , → cryptocurrency,.
内容的提问来源于stack exchange,提问作者Gabriel

