You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python Requests和BS4抓取含Cookies与javax.faces.ViewState的JSF站点

Scraping JSF Site with Requests & BeautifulSoup (Target: SNIRH Hidroweb)

Got it, let's walk through how to scrape this JSF site using Python's requests and BeautifulSoup libraries. JSF sites are a bit finicky because they rely on session cookies and the javax.faces.ViewState parameter to track user interactions—so we need to handle these correctly at each step. Here's a step-by-step implementation tailored to your use case:

Step 1: Set Up Dependencies & Session

First, make sure you have the required libraries installed:

pip install requests beautifulsoup4

We'll use requests.Session() to persist cookies across requests, which is critical for maintaining the JSF session state.

import requests
from bs4 import BeautifulSoup
import re

# Initialize session and set browser-like headers to avoid being blocked
session = requests.Session()
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36',
    'Accept-Language': 'pt-BR,pt;q=0.9,en;q=0.8'
}

base_url = "http://www.snirh.gov.br/hidroweb/publico/medicoes_historicas_abas.jsf"

Step 2: Fetch Initial Page & Extract ViewState

First, we need to load the initial page to get the session cookies and the initial javax.faces.ViewState (this might be in a hidden input or embedded in a JavaScript CDATA section, depending on the site):

# Get the initial page to establish session and grab ViewState
response = session.get(base_url, headers=headers)
soup = BeautifulSoup(response.content, 'html.parser')

# Extract ViewState (handle both hidden input and CDATA cases)
view_state = None
# First check for hidden input
view_state_input = soup.find('input', {'name': 'javax.faces.ViewState'})
if view_state_input:
    view_state = view_state_input['value']
else:
    # If in CDATA, parse script tags
    script_tags = soup.find_all('script', type='text/javascript')
    for script in script_tags:
        if script.string and 'javax.faces.ViewState' in script.string:
            match = re.search(r'javax.faces.ViewState" value="(.*?)"', script.string)
            if match:
                view_state = match.group(1)
                break

if not view_state:
    raise ValueError("Could not extract javax.faces.ViewState")

Step 3: Switch to "Séries Históricas" Tab

Clicking the "Séries Históricas" tab triggers an AJAX request to switch the view. We need to replicate this request using the ViewState we just extracted:

# Find the tab's ID dynamically (avoids hardcoding IDs that might change)
tab_link = soup.find('a', string='Séries Históricas')
if not tab_link:
    raise ValueError("Could not find 'Séries Históricas' tab")
tab_id = tab_link['id']

# Construct the AJAX request data for switching tabs
switch_payload = {
    'javax.faces.partial.ajax': 'true',
    'javax.faces.source': tab_id,
    'javax.faces.partial.execute': tab_id,
    'javax.faces.partial.render': tab_id.split(':')[0],  # Render the parent container
    tab_id: tab_id,
    'javax.faces.ViewState': view_state
}

# Send the tab switch request
session.post(base_url, data=switch_payload, headers=headers)

# After switching, fetch the updated page to get the new ViewState and form details
updated_response = session.get(base_url, headers=headers)
updated_soup = BeautifulSoup(updated_response.content, 'html.parser')

# Extract the new ViewState for the form submission
new_view_state = updated_soup.find('input', {'name': 'javax.faces.ViewState'})['value']

Step 4: Submit the Station Code & Fetch Data

Now we can fill in the station code and submit the "Consultar" button. Again, we'll dynamically find the form elements to avoid hardcoding IDs:

station_code = "938001"  # Your example station code

# Find the form input for station code and the consult button
code_input = updated_soup.find('input', {'placeholder': 'Código da Estação'})
if not code_input:
    code_input = updated_soup.find('input', {'id': re.compile(r'codigoEstacao')})
code_input_name = code_input['name']

consult_button = updated_soup.find('button', string='Consultar')
if not consult_button:
    consult_button = updated_soup.find('input', {'value': 'Consultar'})
consult_button_name = consult_button['name']

# Construct the form submission payload
query_payload = {
    code_input_name: station_code,
    consult_button_name: 'Consultar',
    'javax.faces.ViewState': new_view_state
}

# Send the query request
result_response = session.post(base_url, data=query_payload, headers=headers)
result_soup = BeautifulSoup(result_response.content, 'html.parser')

Step 5: Parse & Extract Target Data

Finally, we can parse the results (assuming the data is in a table, which is common for this site):

# Find the data table (adjust the selector if needed)
data_table = result_soup.find('table', class_='tabela')
if not data_table:
    print("No data table found—check if the station code is valid or the site structure changed")
else:
    # Extract rows and columns
    rows = data_table.find_all('tr')
    for row in rows[1:]:  # Skip header row
        columns = [col.text.strip() for col in row.find_all('td')]
        print(columns)

Key Notes for JSF Scraping

  • Session Persistence: Always use requests.Session() to keep cookies alive—JSF relies on session state to track your interactions.
  • Dynamic ViewState: The javax.faces.ViewState changes after most POST requests, so you need to re-extract it before each subsequent submission.
  • Avoid Hardcoded IDs: JSF generates dynamic component IDs (like j_idt32) that can change when the site updates. Use text or attribute patterns to find elements dynamically.
  • Browser Emulation: Set a realistic User-Agent and accept-language header to avoid being flagged as a bot.
  • Debug with DevTools: If something breaks, use your browser's DevTools (Network tab) to inspect the actual requests sent when you click buttons or tabs—replicate those headers and payloads exactly.

内容的提问来源于stack exchange,提问作者pedro regis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 07:40:46