R语言含中间页面的网页爬取问题求助
Hey there! Let's tackle this web scraping hurdle together. That intermediate page (http://tempest.wellesley.edu/~btjaden/cgi-bin/processRequest2.cgi) is likely a session-based redirect or confirmation step—meaning your script needs to maintain persistent cookies and follow the request flow properly to access the results text file. Here's a step-by-step solution using R's httr and rvest packages, which handle sessions and HTML parsing seamlessly:
Step 1: Set up your environment
First, make sure you have the necessary packages installed and loaded:
# Install packages if you haven't already if (!require("httr")) install.packages("httr") if (!require("rvest")) install.packages("rvest") library(httr) library(rvest)
Step 2: Create a persistent session
The key here is to use a session object to carry over cookies and form data between requests—this mimics how a browser would handle the form submission and intermediate page:
# Start a session on the original form page (replace with your actual form URL) my_session <- session("http://tempest.wellesley.edu/~btjaden/")
Step 3: Submit your form with specific values
Define all the required form parameters (use your browser's DevTools to inspect the form fields, including hidden ones!) and submit them via the session:
# Replace these with your actual form field names and values form_params <- list( query_term = "your_specific_value", another_field = "another_required_value", # Add every field the form expects—check DevTools > Network tab for details ) # Submit the form to the intermediate page (or the original form's action URL) form_response <- session_post( my_session, url = "http://tempest.wellesley.edu/~btjaden/cgi-bin/processRequest2.cgi", body = form_params, encode = "form" )
Step 4: Extract the "View results as text file" link
Parse the intermediate page's HTML to grab the results text file link:
# Convert the response to an HTML object results_page <- read_html(form_response) # Extract the link (adjust the selector if needed—use DevTools to get the right element) text_file_link <- results_page %>% html_element("a:contains('View results as text file')") %>% html_attr("href") # If the link is relative, convert it to an absolute URL text_file_link <- paste0("http://tempest.wellesley.edu/~btjaden/", text_file_link)
Step 5: Fetch and parse the text file into a data.frame
Use the session to fetch the text file content, then convert it to a data.frame:
# Get the text file content text_content <- session_get(my_session, url = text_file_link) %>% content(as = "text") # Parse into a data.frame (adjust sep/header based on the actual text format) results_df <- read.table( text = text_content, header = TRUE, sep = "\t", # Common for tab-separated text files—change to "," for CSV if needed stringsAsFactors = FALSE ) # Check your results! head(results_df)
Key Notes to Troubleshoot:
- Double-check form parameters: Use your browser's Network tab to capture exactly what fields the form sends—don't miss hidden inputs, as many CGI scripts require them.
- Adjust selectors: If
a:contains()doesn't find the link, use a more specific selector (e.g., byidorclassif the link has one). - Handle redirects: If the intermediate page auto-redirects,
session_postwill usually follow it automatically—you can skip the link extraction step and parse the redirected page directly.
内容的提问来源于stack exchange,提问作者tatty_dk

