使用rvest提取Morningstar页面Investment Objective内容的分页问题
Got it, let's fix this problem for you. The core issue here is that the SEC prospectus embedded in the iframe is paginated—your current code only pulls the first page, which doesn't include the "Investment Objective" section. Here's how to adjust your R code to target the right page or grab the full document directly:
Step 1: Retrieve the iframe's source URL
First, we'll confirm the actual SEC page URL embedded in the iframe, just like your original code did:
library(rvest) library(httr) # For handling relative URLs library(stringr) # Original Morningstar link linke <- 'http://quote.morningstar.com/etf-filing/Prospectus/2018/1/18/t.aspx?t=SPY&ft=497&d=0833554effb2f4d14d1f23a561738303' # Extract the iframe's source URL iframe_src <- read_html(linke) %>% html_node("iframe.sec_frame") %>% html_attr("src")
Step 2: Option 1: Target the Second Page Directly
The SEC prospectus uses pagination with page parameters in the URL. We'll find the link to the second page, convert it to a full URL, then extract the target content:
# Load the initial SEC page sec_page <- read_html(iframe_src) # Find the link to the second page (filter for links containing "page=2") page_2_link <- sec_page %>% html_nodes("a") %>% html_attr("href") %>% str_subset("page=2") %>% head(1) # Convert relative link to full URL second_page_url <- url_absolute(page_2_link, iframe_src) # Load the second page and extract Investment Objective second_page <- read_html(second_page_url) investment_objective <- second_page %>% html_node(xpath=".//div[contains(., 'Investment Objective')]") %>% html_text() # Print the result cat(investment_objective)
Step 3: Option 2: Grab the Full Prospectus Text (More Reliable)
If pagination links change, a better fallback is to pull the complete plain-text version of the prospectus from the SEC's site:
# Find the link to the full plain-text document full_text_link <- sec_page %>% html_nodes("a") %>% html_attr("href") %>% str_subset("\\.txt$") %>% head(1) # Convert to full URL and load the text full_text_url <- url_absolute(full_text_link, iframe_src) full_prospectus <- read_lines(full_text_url) # Extract the Investment Objective section (adjust boundaries as needed) objective_start <- str_which(full_prospectus, "Investment Objective") objective_end <- str_which(full_prospectus, "Investment Strategies") - 1 investment_objective_full <- full_prospectus[objective_start:objective_end] # Print the cleaned section cat(paste(investment_objective_full, collapse = "\n"))
Why Your Original Code Failed
The iframe loads a paginated SEC document, and the "Investment Objective" section lives on the second page (not the first). Your code only fetched the initial page, so it couldn't find the target content.
内容的提问来源于stack exchange,提问作者Talha Naushad

