在R语言中使用正则表达式提取HTML字符串中的姓名与职位
Got it, let's tackle this problem step by step. Since you've already converted your HTML content into character vectors, regular expressions are perfect for pulling out the names and positions here—no need for grepl filtering anymore.
Step 1: Regex Patterns for Extraction
First, let's define the regex patterns tailored to your specific HTML structure:
For Names: We need to capture text between
<h3 class="personName">and</h3>. Using positive lookbehind/lookahead ensures we only grab the content, not the surrounding tags:(?<=<h3 class="personName">).*?(?=</h3>)The
.*?is a non-greedy match—it stops at the first closing</h3>tag, which prevents accidental matching of extra content if your vectors ever have more tags later.For Positions: Similarly, capture text between
<li>and</li>with this pattern:(?<=<li>).*?(?=</li>)
Step 2: Implement in R (with 2 approaches)
Since you mentioned grepl, I assume you're working in R. Here are two straightforward ways to apply these regex patterns:
Option 1: Using stringr (cleaner syntax)
The stringr package makes regex extraction super intuitive:
# Your sample character vector html_vectors <- c( '<h3 class="personName">Whitney Alicia Zimmerman</h3> <li>Assistant Teaching Professor</li>', '<h3 class="personName">John Michael Doe</h3> <li>Senior Research Fellow</li>' ) library(stringr) # Extract names and positions extracted_names <- str_extract(html_vectors, '(?<=<h3 class="personName">).*?(?=</h3>)') extracted_positions <- str_extract(html_vectors, '(?<=<li>).*?(?=</li>)') # Combine into a tidy data frame result <- data.frame( Full_Name = extracted_names, Position = extracted_positions ) print(result)
Option 2: Base R (no extra packages needed)
If you prefer sticking to base R, use regexpr with regmatches:
# Sample vector (same as above) html_vectors <- c( '<h3 class="personName">Whitney Alicia Zimmerman</h3> <li>Assistant Teaching Professor</li>', '<h3 class="personName">John Michael Doe</h3> <li>Senior Research Fellow</li>' ) # Extract names name_matches <- regexpr('(?<=<h3 class="personName">).*?(?=</h3>)', html_vectors, perl = TRUE) extracted_names <- regmatches(html_vectors, name_matches) # Extract positions position_matches <- regexpr('(?<=<li>).*?(?=</li>)', html_vectors, perl = TRUE) extracted_positions <- regmatches(html_vectors, position_matches) # Combine into data frame result <- data.frame( Full_Name = extracted_names, Position = extracted_positions ) print(result)
Step 3: Make it more robust (optional)
If your HTML tags might have extra attributes (like id or additional classes), adjust the name pattern to handle that:
(?<=<h3[^>]+class="personName"[^>]+>).*?(?=</h3>)
The [^>]+ matches any characters except the closing >, so it works even if the <h3> tag has other properties.
内容的提问来源于stack exchange,提问作者Zack Vlliet

