You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用rvest获取特定标签后的#text内容?网页台词抓取求助

How to Scrape Both Character Names and Lines from the Avatar Transcript Page

Hey there! I see you've already nailed grabbing character names using <b> tags, but pulling their corresponding lines (which live in sibling text nodes) is a common rvest gotcha. Let's fix that together.

The Core Issue

On that transcript page, each bolded character name (<b>) is immediately followed by a raw text node (#text) containing their line. Since html_nodes("b") only targets the bolded elements, you need to explicitly grab the sibling text node right after each <b> tag to get the lines.

Solution Code

Here are two reliable ways to pair characters with their lines into a tidy data frame:

Method 1: Iterate Through Nodes (Most Precise)

library(rvest)
library(dplyr)

# Load the target page
url1 <- "http://atla.avatarspirit.net/transcripts.php?num=120"
webpage <- read_html(url1)

# Scrape paired characters and lines
transcript <- webpage %>%
  html_nodes("b") %>%
  map_dfr(function(node) {
    # Extract the character name from the <b> tag
    character <- node %>% html_text(trim = TRUE)
    # Grab the first text node immediately after the <b> tag
    line <- node %>% 
      html_node(xpath = "following-sibling::text()[1]") %>% 
      html_text(trim = TRUE)
    # Return as a row in a data frame
    tibble(character, line)
  })

# Preview the results
head(transcript)

Method 2: Bulk Node Selection (Concise Alternative)

If you prefer a shorter approach, you can pull all text siblings first and pair them with character names directly:

transcript_bulk <- tibble(
  character = webpage %>% html_nodes("b") %>% html_text(trim = TRUE),
  line = webpage %>% html_nodes(xpath = "//b/following-sibling::text()[1]") %>% html_text(trim = TRUE)
)

Why This Works

  • following-sibling::text()[1] is an XPath selector that targets the first text node immediately after each <b> tag—exactly where the character's line is stored.
  • trim = TRUE cleans up extra whitespace and line breaks that clutter raw HTML text nodes.

Quick Cleanup Tip

If you notice empty lines in the results, filter them out with:

transcript_clean <- transcript %>% filter(line != "")

内容的提问来源于stack exchange,提问作者Kira Tebbe

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:47:41