在R语言中如何用正则表达式匹配斜体与下划线格式内容
Hey there! Let's tackle your regex questions for R step by step.
Your first goal is to match the 1st and 3rd strings in your vector, which contain <em> opening and closing tags.
The regex pattern targets any string that includes at least one pair of <em> and </em> tags (with any content in between, plus any text before or after the tags). Here's the pattern and how to use it in R:
# Define the regex pattern em_pattern <- "^.*<em>.*</em>.*$" # Your original vector ex <- c("what <em>did</em> you do?", "i don鈥檛 know", "<em>Pe</em>ter is angry.") # Check matches with grepl() grepl(em_pattern, ex) # Output: TRUE FALSE TRUE
Breakdown of the pattern:
^anchors the match to the start of the string.*matches any character (except newlines) zero or more times<em>and</em>match the literal HTML tags exactly$anchors the match to the end of the string
If the italicized content uses underscores instead (e.g., what _did_ you do? and _Pe_ter is angry.), we adjust the regex to target underscore-wrapped content.
For precise matching (ensuring there's actual text between the underscores), use this pattern:
# Define the underscore regex pattern underscore_pattern <- "^.*_[^_]+_.*$" # Updated vector with underscore formatting ex_underscore <- c("what _did_ you do?", "i don鈥檛 know", "_Pe_ter is angry.") # Check matches grepl(underscore_pattern, ex_underscore) # Output: TRUE FALSE TRUE
Breakdown of the pattern:
_matches the literal underscore character[^_]+matches one or more characters that are NOT underscores (prevents matching empty__pairs)^.*and.*$handle any text before or after the underscore-wrapped content
If you don't need to worry about empty underscore pairs, you could simplify the pattern to ^.*_.*_.*$—it works for your specific example, but the [^_]+ version is more robust for real-world use.
内容的提问来源于stack exchange,提问作者Chris Ruehlemann

