You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用rvest提取Morningstar页面Investment Objective内容的分页问题

Fixing the rvest Issue: Grabbing the Second Page of the Embedded Prospectus

Got it, let's fix this problem for you. The core issue here is that the SEC prospectus embedded in the iframe is paginated—your current code only pulls the first page, which doesn't include the "Investment Objective" section. Here's how to adjust your R code to target the right page or grab the full document directly:

Step 1: Retrieve the iframe's source URL

First, we'll confirm the actual SEC page URL embedded in the iframe, just like your original code did:

library(rvest)
library(httr) # For handling relative URLs
library(stringr)

# Original Morningstar link
linke <- 'http://quote.morningstar.com/etf-filing/Prospectus/2018/1/18/t.aspx?t=SPY&ft=497&d=0833554effb2f4d14d1f23a561738303'

# Extract the iframe's source URL
iframe_src <- read_html(linke) %>% 
  html_node("iframe.sec_frame") %>% 
  html_attr("src")

Step 2: Option 1: Target the Second Page Directly

The SEC prospectus uses pagination with page parameters in the URL. We'll find the link to the second page, convert it to a full URL, then extract the target content:

# Load the initial SEC page
sec_page <- read_html(iframe_src)

# Find the link to the second page (filter for links containing "page=2")
page_2_link <- sec_page %>% 
  html_nodes("a") %>% 
  html_attr("href") %>% 
  str_subset("page=2") %>% 
  head(1)

# Convert relative link to full URL
second_page_url <- url_absolute(page_2_link, iframe_src)

# Load the second page and extract Investment Objective
second_page <- read_html(second_page_url)
investment_objective <- second_page %>% 
  html_node(xpath=".//div[contains(., 'Investment Objective')]") %>% 
  html_text()

# Print the result
cat(investment_objective)

Step 3: Option 2: Grab the Full Prospectus Text (More Reliable)

If pagination links change, a better fallback is to pull the complete plain-text version of the prospectus from the SEC's site:

# Find the link to the full plain-text document
full_text_link <- sec_page %>% 
  html_nodes("a") %>% 
  html_attr("href") %>% 
  str_subset("\\.txt$") %>% 
  head(1)

# Convert to full URL and load the text
full_text_url <- url_absolute(full_text_link, iframe_src)
full_prospectus <- read_lines(full_text_url)

# Extract the Investment Objective section (adjust boundaries as needed)
objective_start <- str_which(full_prospectus, "Investment Objective")
objective_end <- str_which(full_prospectus, "Investment Strategies") - 1
investment_objective_full <- full_prospectus[objective_start:objective_end]

# Print the cleaned section
cat(paste(investment_objective_full, collapse = "\n"))

Why Your Original Code Failed

The iframe loads a paginated SEC document, and the "Investment Objective" section lives on the second page (not the first). Your code only fetched the initial page, so it couldn't find the target content.

内容的提问来源于stack exchange,提问作者Talha Naushad

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 04:27:47