You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用R的rvest工具无CSS选择器抓取爵士专辑人员信息的方法

How to Scrape Unmarked Personnel Text with rvest for Blue Note Catalog Pages

Got it, let's work through this problem! When you're dealing with text that doesn't have a dedicated CSS selector (like those personnel lists starting with "Sonny Rollins, tenor sax..."), XPath and text filtering with regex become your best tools—they're far more flexible than CSS for targeting unstructured text nodes. Here are three reliable approaches using rvest:

1. Target Parent Containers, Then Filter Text

First, identify the parent element that wraps each album's full information (like a table row or div). Extract all text from these containers, then use regex to isolate the personnel lines.

library(rvest)
library(stringr)

url <- "https://www.jazzdisco.org/blue-note-records/catalog-4000-series/"
page <- read_html(url)

# Grab each album's container (adjust the selector to match the page's actual structure)
album_blocks <- page %>% html_elements("tr[valign='top']")

# Extract personnel from each block
personnel_list <- lapply(album_blocks, function(block) {
  # Split container text into clean lines
  all_lines <- block %>% html_text2() %>% str_split("\n") %>% unlist()
  
  # Match lines that follow the "Name, Instrument" pattern
  personnel_line <- all_lines[grepl("^[A-Z][a-z]+ [A-Z][a-z]+, ", all_lines)]
  
  # Optional: refine regex to target common jazz instruments for precision
  # personnel_line <- all_lines[grepl(", (tenor sax|piano|drums|trumpet|bass)", all_lines)]
  
  trimws(personnel_line)
})

# Filter out empty entries (for rows without personnel)
personnel_list <- Filter(length, personnel_list)

2. Use XPath to Directly Target Text Nodes

XPath lets you select text nodes that sit next to elements with known selectors. For example, if personnel text comes right after an album title in an <h3> tag:

# Target text nodes that follow an album title and contain the ", " name-instrument separator
personnel_nodes <- page %>% html_elements(xpath = "//td/h3/following-sibling::text()[contains(., ', ')]")

# Clean up the text
personnel_text <- personnel_nodes %>% html_text() %>% trimws()

3. Extract from Parent Elements with Context

If personnel text shares a parent with other marked elements (like catalog numbers or album titles), extract the full parent text and use regex to pull out the personnel section:

album_rows <- page %>% html_elements("table tr")

personnel_list <- lapply(album_rows, function(row) {
  full_text <- row %>% html_text2()
  
  # Extract everything after the catalog number/title (adjust regex to match page structure)
  personnel <- str_extract(full_text, "(?<=\\d{4} - .+\\n).+")
  
  trimws(personnel)
})

Pro Tips:

  • Always use your browser's developer tools to inspect the page structure first—identify consistent parent elements or adjacent marked nodes to anchor your selectors.
  • Test regex patterns incrementally: print out raw text from a single album block first to see how the personnel line is formatted, then tweak your regex to match.
  • Use html_text2() instead of html_text()—it handles line breaks and whitespace more cleanly, making it easier to split and filter text.

内容的提问来源于stack exchange,提问作者Ben

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 06:52:29