You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

R语言含中间页面的网页爬取问题求助

Hey there! Let's tackle this web scraping hurdle together. That intermediate page (http://tempest.wellesley.edu/~btjaden/cgi-bin/processRequest2.cgi) is likely a session-based redirect or confirmation step—meaning your script needs to maintain persistent cookies and follow the request flow properly to access the results text file. Here's a step-by-step solution using R's httr and rvest packages, which handle sessions and HTML parsing seamlessly:

Step 1: Set up your environment

First, make sure you have the necessary packages installed and loaded:

# Install packages if you haven't already
if (!require("httr")) install.packages("httr")
if (!require("rvest")) install.packages("rvest")

library(httr)
library(rvest)

Step 2: Create a persistent session

The key here is to use a session object to carry over cookies and form data between requests—this mimics how a browser would handle the form submission and intermediate page:

# Start a session on the original form page (replace with your actual form URL)
my_session <- session("http://tempest.wellesley.edu/~btjaden/")

Step 3: Submit your form with specific values

Define all the required form parameters (use your browser's DevTools to inspect the form fields, including hidden ones!) and submit them via the session:

# Replace these with your actual form field names and values
form_params <- list(
  query_term = "your_specific_value",
  another_field = "another_required_value",
  # Add every field the form expects—check DevTools > Network tab for details
)

# Submit the form to the intermediate page (or the original form's action URL)
form_response <- session_post(
  my_session,
  url = "http://tempest.wellesley.edu/~btjaden/cgi-bin/processRequest2.cgi",
  body = form_params,
  encode = "form"
)

Parse the intermediate page's HTML to grab the results text file link:

# Convert the response to an HTML object
results_page <- read_html(form_response)

# Extract the link (adjust the selector if needed—use DevTools to get the right element)
text_file_link <- results_page %>% 
  html_element("a:contains('View results as text file')") %>% 
  html_attr("href")

# If the link is relative, convert it to an absolute URL
text_file_link <- paste0("http://tempest.wellesley.edu/~btjaden/", text_file_link)

Step 5: Fetch and parse the text file into a data.frame

Use the session to fetch the text file content, then convert it to a data.frame:

# Get the text file content
text_content <- session_get(my_session, url = text_file_link) %>% 
  content(as = "text")

# Parse into a data.frame (adjust sep/header based on the actual text format)
results_df <- read.table(
  text = text_content,
  header = TRUE,
  sep = "\t", # Common for tab-separated text files—change to "," for CSV if needed
  stringsAsFactors = FALSE
)

# Check your results!
head(results_df)

Key Notes to Troubleshoot:

  • Double-check form parameters: Use your browser's Network tab to capture exactly what fields the form sends—don't miss hidden inputs, as many CGI scripts require them.
  • Adjust selectors: If a:contains() doesn't find the link, use a more specific selector (e.g., by id or class if the link has one).
  • Handle redirects: If the intermediate page auto-redirects, session_post will usually follow it automatically—you can skip the link extraction step and parse the redirected page directly.

内容的提问来源于stack exchange,提问作者tatty_dk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 10:28:03