如何用rvest获取特定标签后的#text内容?网页台词抓取求助
Hey there! I see you've already nailed grabbing character names using <b> tags, but pulling their corresponding lines (which live in sibling text nodes) is a common rvest gotcha. Let's fix that together.
The Core Issue
On that transcript page, each bolded character name (<b>) is immediately followed by a raw text node (#text) containing their line. Since html_nodes("b") only targets the bolded elements, you need to explicitly grab the sibling text node right after each <b> tag to get the lines.
Solution Code
Here are two reliable ways to pair characters with their lines into a tidy data frame:
Method 1: Iterate Through Nodes (Most Precise)
library(rvest) library(dplyr) # Load the target page url1 <- "http://atla.avatarspirit.net/transcripts.php?num=120" webpage <- read_html(url1) # Scrape paired characters and lines transcript <- webpage %>% html_nodes("b") %>% map_dfr(function(node) { # Extract the character name from the <b> tag character <- node %>% html_text(trim = TRUE) # Grab the first text node immediately after the <b> tag line <- node %>% html_node(xpath = "following-sibling::text()[1]") %>% html_text(trim = TRUE) # Return as a row in a data frame tibble(character, line) }) # Preview the results head(transcript)
Method 2: Bulk Node Selection (Concise Alternative)
If you prefer a shorter approach, you can pull all text siblings first and pair them with character names directly:
transcript_bulk <- tibble( character = webpage %>% html_nodes("b") %>% html_text(trim = TRUE), line = webpage %>% html_nodes(xpath = "//b/following-sibling::text()[1]") %>% html_text(trim = TRUE) )
Why This Works
following-sibling::text()[1]is an XPath selector that targets the first text node immediately after each<b>tag—exactly where the character's line is stored.trim = TRUEcleans up extra whitespace and line breaks that clutter raw HTML text nodes.
Quick Cleanup Tip
If you notice empty lines in the results, filter them out with:
transcript_clean <- transcript %>% filter(line != "")
内容的提问来源于stack exchange,提问作者Kira Tebbe

