You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从GET请求返回的HTML响应中提取NASA页面的表格数据

Hey Pedro, let’s work through how to pull that table data you need from the NASA EOS SSE RETScreen page. Since the response is a fully rendered HTML page, you’ve got two reliable approaches here—let’s break them down:

Option 1: Tap into the underlying API (the cleanest, fastest method)

Most form-based data pages like this don’t generate the HTML directly on the server—they fetch raw data via a backend API, then render it into the table in your browser. Skipping the HTML parsing entirely by hitting that API directly will save you a lot of hassle. Here’s how to find it:

  • Open your browser’s DevTools (press F12) and switch to the Network tab.
  • Submit the form with your target latitude/longitude like you normally would.
  • Filter the requests by Fetch/XHR to narrow down to data-focused calls. Look for requests that fire right after you submit the form—this is likely the one returning your raw data.
  • Check the request’s payload (you’ll see your lat/long and any other parameters sent) and the response. It might be structured as JSON, CSV, or XML—way easier to parse than HTML.
  • Once you’ve identified the API endpoint and required parameters, use a tool like Python’s requests library to send the same payload programmatically and get clean, structured data directly.
Option 2: Use browser automation to extract rendered table data

If you can’t track down the API (or need to fully simulate the user flow), browser automation tools like Selenium or Playwright will load the page, run all the JavaScript to render the table, and let you extract the data. Here’s a step-by-step with a Python/Selenium example:

  • First, set up Selenium with your preferred browser driver (Chrome, Firefox, etc.).
  • Navigate to the RETScreen page, automate filling in the lat/long fields, and submit the form.
  • Wait explicitly for the table to load (don’t use hard sleeps—use Selenium’s WebDriverWait to ensure the table is present before trying to extract data).
  • Target the table and loop through its rows/cells to pull your specified columns.

Here’s a quick code snippet to get you started:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# Initialize the Chrome driver
driver = webdriver.Chrome()
driver.get("https://eosweb.larc.nasa.gov/sse/RETScreen/")

# Fill in latitude and longitude (adjust selectors if the page uses different names/IDs)
lat_field = driver.find_element(By.NAME, "lat")
lon_field = driver.find_element(By.NAME, "lon")
lat_field.send_keys("37.7749")  # Example: San Francisco latitude
lon_field.send_keys("-122.4194") # Example: San Francisco longitude

# Submit the form (update the selector to match the actual submit button)
submit_button = driver.find_element(By.CSS_SELECTOR, "input[type='submit'][value='Submit']")
submit_button.click()

# Wait for the target table to load (adjust the CSS selector to match your table's class/ID)
target_table = WebDriverWait(driver, 15).until(
    EC.presence_of_element_located((By.CSS_SELECTOR, "table.sse-data-table"))
)

# Define which columns you want to extract (0-indexed)
target_column_indices = [2, 4]  # Example: 3rd and 5th columns

# Extract rows (skip the header row with [1:])
rows = target_table.find_elements(By.TAG_NAME, "tr")
for row in rows[1:]:
    cells = row.find_elements(By.TAG_NAME, "td")
    # Only process rows that have enough cells to cover your target columns
    if len(cells) > max(target_column_indices):
        extracted_values = [cells[idx].text for idx in target_column_indices]
        print(extracted_values)

# Clean up
driver.quit()
Quick Pro Tips
  • Always inspect the page elements first (right-click > Inspect) to get accurate CSS selectors or XPath for inputs, buttons, and tables—unique IDs or classes will make targeting way more reliable.
  • If using the API approach, make sure to copy any required request headers (like User-Agent) from the DevTools request to avoid being blocked by NASA’s servers.
  • For large datasets, the API method is always preferable—it’s faster, uses less bandwidth, and avoids the fragility of scraping HTML (which can break if the page’s design changes).

内容的提问来源于stack exchange,提问作者Pedro Miguel Pimienta Morales

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.20 11:55:20