如何用R提取搜索结果的entrylist_link并下载对应文章?
Hey there! Let's work through getting those article links extracted and their content downloaded—you're already halfway there with the title scraping. Here's how to tackle this:
Step 1: Fix the Link Extraction
First, the issue with grabbing the href attributes is probably just a matter of targeting the right CSS selector. Looking at the Süddeutsche Zeitung's search results, the article links are wrapped in <a> tags with the class entrylist__link (double-check with your browser's dev tools if this varies slightly).
Here's how to extract those links properly, including fixing relative paths to full URLs:
library(rvest) library(stringr) library(purrr) library(tibble) # Your original search URL url_parsed1 <- read_html("http://www.sueddeutsche.de/news?search=Fl%C3%BCchtlinge&sort=date&dep%5B%5D=politik&typ%5B%5D=article&sys%5B%5D=sz&catsz%5B%5D=alles&time=2015-01-01T00%3A00%2F2015-12-31T23%3A59&startDate=01.01.2015&endDate=31.01.2015") # Extract the <a> tags that hold the article links article_link_nodes <- html_nodes(url_parsed1, css = ".entrylist__link") # Grab the href attribute from each node article_links <- html_attr(article_link_nodes, "href") # Convert relative URLs to full absolute URLs (critical for accessing the articles) article_links_full <- url_absolute(article_links, base_url = "http://www.sueddeutsche.de") # Check the first few links to confirm they're correct head(article_links_full)
Step 2: Scrape Content from Each Article
Now that you have full links, create a helper function to pull the title and body from each article, then loop through all links. I've added error handling and a small delay to avoid triggering anti-scraping measures (always be respectful of a website's terms of service!):
# Helper function to safely scrape a single article scrape_single_article <- function(link) { # Add a 2-second delay between requests to avoid being blocked Sys.sleep(2) # Use tryCatch to handle broken links or unexpected page structures tryCatch({ article_page <- read_html(link) # Extract article title (adjust selector if needed for your target pages) article_title <- html_text(html_nodes(article_page, css = ".sz-article-header__title")) %>% str_replace_all("\\n|\\t|\\r", "") %>% str_trim() # Extract article body (again, tweak the CSS selector if the page structure changes) article_body <- html_text(html_nodes(article_page, css = ".sz-article-body")) %>% str_replace_all("\\n|\\t|\\r", " ") %>% str_trim() # Return a dataframe row with the content tibble(title = article_title, body = article_body, link = link, error = NA) }, error = function(e) { # Return a row with error details if something goes wrong tibble(title = NA, body = NA, link = link, error = e$message) }) } # Scrape all articles and combine into a single dataframe all_articles <- map_dfr(article_links_full, scrape_single_article) # View the first few results head(all_articles)
Quick Troubleshooting Tips
- CSS Selector Adjustments: If the link or content selectors don't work, right-click the element in your browser and use "Inspect" to find the correct class or tag. Update the
cssparameter inhtml_nodes()accordingly. - Anti-Scraping Mitigation: If you get 403/500 errors, increase the
Sys.sleep()delay or add a user-agent header withhttr::user_agent()when reading pages (e.g.,read_html(link, user_agent("Mozilla/5.0"))). - Data Storage: Save your results to a CSV for later use with
write_csv(all_articles, "sueddeutsche_refugee_articles_jan2015.csv").
内容的提问来源于stack exchange,提问作者matt

