You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用R提取搜索结果的entrylist_link并下载对应文章?

Hey there! Let's work through getting those article links extracted and their content downloaded—you're already halfway there with the title scraping. Here's how to tackle this:

First, the issue with grabbing the href attributes is probably just a matter of targeting the right CSS selector. Looking at the Süddeutsche Zeitung's search results, the article links are wrapped in <a> tags with the class entrylist__link (double-check with your browser's dev tools if this varies slightly).

Here's how to extract those links properly, including fixing relative paths to full URLs:

library(rvest)
library(stringr)
library(purrr)
library(tibble)

# Your original search URL
url_parsed1 <- read_html("http://www.sueddeutsche.de/news?search=Fl%C3%BCchtlinge&amp;sort=date&amp;dep%5B%5D=politik&amp;typ%5B%5D=article&amp;sys%5B%5D=sz&amp;catsz%5B%5D=alles&amp;time=2015-01-01T00%3A00%2F2015-12-31T23%3A59&amp;startDate=01.01.2015&amp;endDate=31.01.2015")

# Extract the <a> tags that hold the article links
article_link_nodes <- html_nodes(url_parsed1, css = ".entrylist__link")
# Grab the href attribute from each node
article_links <- html_attr(article_link_nodes, "href")

# Convert relative URLs to full absolute URLs (critical for accessing the articles)
article_links_full <- url_absolute(article_links, base_url = "http://www.sueddeutsche.de")

# Check the first few links to confirm they're correct
head(article_links_full)

Step 2: Scrape Content from Each Article

Now that you have full links, create a helper function to pull the title and body from each article, then loop through all links. I've added error handling and a small delay to avoid triggering anti-scraping measures (always be respectful of a website's terms of service!):

# Helper function to safely scrape a single article
scrape_single_article <- function(link) {
  # Add a 2-second delay between requests to avoid being blocked
  Sys.sleep(2)
  
  # Use tryCatch to handle broken links or unexpected page structures
  tryCatch({
    article_page <- read_html(link)
    
    # Extract article title (adjust selector if needed for your target pages)
    article_title <- html_text(html_nodes(article_page, css = ".sz-article-header__title")) %>%
      str_replace_all("\\n|\\t|\\r", "") %>%
      str_trim()
    
    # Extract article body (again, tweak the CSS selector if the page structure changes)
    article_body <- html_text(html_nodes(article_page, css = ".sz-article-body")) %>%
      str_replace_all("\\n|\\t|\\r", " ") %>%
      str_trim()
    
    # Return a dataframe row with the content
    tibble(title = article_title, body = article_body, link = link, error = NA)
  }, error = function(e) {
    # Return a row with error details if something goes wrong
    tibble(title = NA, body = NA, link = link, error = e$message)
  })
}

# Scrape all articles and combine into a single dataframe
all_articles <- map_dfr(article_links_full, scrape_single_article)

# View the first few results
head(all_articles)

Quick Troubleshooting Tips

  • CSS Selector Adjustments: If the link or content selectors don't work, right-click the element in your browser and use "Inspect" to find the correct class or tag. Update the css parameter in html_nodes() accordingly.
  • Anti-Scraping Mitigation: If you get 403/500 errors, increase the Sys.sleep() delay or add a user-agent header with httr::user_agent() when reading pages (e.g., read_html(link, user_agent("Mozilla/5.0"))).
  • Data Storage: Save your results to a CSV for later use with write_csv(all_articles, "sueddeutsche_refugee_articles_jan2015.csv").

内容的提问来源于stack exchange,提问作者matt

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 06:54:19