如何用Python Requests和BS4抓取含Cookies与javax.faces.ViewState的JSF站点
Got it, let's walk through how to scrape this JSF site using Python's requests and BeautifulSoup libraries. JSF sites are a bit finicky because they rely on session cookies and the javax.faces.ViewState parameter to track user interactions—so we need to handle these correctly at each step. Here's a step-by-step implementation tailored to your use case:
Step 1: Set Up Dependencies & Session
First, make sure you have the required libraries installed:
pip install requests beautifulsoup4
We'll use requests.Session() to persist cookies across requests, which is critical for maintaining the JSF session state.
import requests from bs4 import BeautifulSoup import re # Initialize session and set browser-like headers to avoid being blocked session = requests.Session() headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36', 'Accept-Language': 'pt-BR,pt;q=0.9,en;q=0.8' } base_url = "http://www.snirh.gov.br/hidroweb/publico/medicoes_historicas_abas.jsf"
Step 2: Fetch Initial Page & Extract ViewState
First, we need to load the initial page to get the session cookies and the initial javax.faces.ViewState (this might be in a hidden input or embedded in a JavaScript CDATA section, depending on the site):
# Get the initial page to establish session and grab ViewState response = session.get(base_url, headers=headers) soup = BeautifulSoup(response.content, 'html.parser') # Extract ViewState (handle both hidden input and CDATA cases) view_state = None # First check for hidden input view_state_input = soup.find('input', {'name': 'javax.faces.ViewState'}) if view_state_input: view_state = view_state_input['value'] else: # If in CDATA, parse script tags script_tags = soup.find_all('script', type='text/javascript') for script in script_tags: if script.string and 'javax.faces.ViewState' in script.string: match = re.search(r'javax.faces.ViewState" value="(.*?)"', script.string) if match: view_state = match.group(1) break if not view_state: raise ValueError("Could not extract javax.faces.ViewState")
Step 3: Switch to "Séries Históricas" Tab
Clicking the "Séries Históricas" tab triggers an AJAX request to switch the view. We need to replicate this request using the ViewState we just extracted:
# Find the tab's ID dynamically (avoids hardcoding IDs that might change) tab_link = soup.find('a', string='Séries Históricas') if not tab_link: raise ValueError("Could not find 'Séries Históricas' tab") tab_id = tab_link['id'] # Construct the AJAX request data for switching tabs switch_payload = { 'javax.faces.partial.ajax': 'true', 'javax.faces.source': tab_id, 'javax.faces.partial.execute': tab_id, 'javax.faces.partial.render': tab_id.split(':')[0], # Render the parent container tab_id: tab_id, 'javax.faces.ViewState': view_state } # Send the tab switch request session.post(base_url, data=switch_payload, headers=headers) # After switching, fetch the updated page to get the new ViewState and form details updated_response = session.get(base_url, headers=headers) updated_soup = BeautifulSoup(updated_response.content, 'html.parser') # Extract the new ViewState for the form submission new_view_state = updated_soup.find('input', {'name': 'javax.faces.ViewState'})['value']
Step 4: Submit the Station Code & Fetch Data
Now we can fill in the station code and submit the "Consultar" button. Again, we'll dynamically find the form elements to avoid hardcoding IDs:
station_code = "938001" # Your example station code # Find the form input for station code and the consult button code_input = updated_soup.find('input', {'placeholder': 'Código da Estação'}) if not code_input: code_input = updated_soup.find('input', {'id': re.compile(r'codigoEstacao')}) code_input_name = code_input['name'] consult_button = updated_soup.find('button', string='Consultar') if not consult_button: consult_button = updated_soup.find('input', {'value': 'Consultar'}) consult_button_name = consult_button['name'] # Construct the form submission payload query_payload = { code_input_name: station_code, consult_button_name: 'Consultar', 'javax.faces.ViewState': new_view_state } # Send the query request result_response = session.post(base_url, data=query_payload, headers=headers) result_soup = BeautifulSoup(result_response.content, 'html.parser')
Step 5: Parse & Extract Target Data
Finally, we can parse the results (assuming the data is in a table, which is common for this site):
# Find the data table (adjust the selector if needed) data_table = result_soup.find('table', class_='tabela') if not data_table: print("No data table found—check if the station code is valid or the site structure changed") else: # Extract rows and columns rows = data_table.find_all('tr') for row in rows[1:]: # Skip header row columns = [col.text.strip() for col in row.find_all('td')] print(columns)
Key Notes for JSF Scraping
- Session Persistence: Always use
requests.Session()to keep cookies alive—JSF relies on session state to track your interactions. - Dynamic ViewState: The
javax.faces.ViewStatechanges after most POST requests, so you need to re-extract it before each subsequent submission. - Avoid Hardcoded IDs: JSF generates dynamic component IDs (like
j_idt32) that can change when the site updates. Use text or attribute patterns to find elements dynamically. - Browser Emulation: Set a realistic
User-Agentand accept-language header to avoid being flagged as a bot. - Debug with DevTools: If something breaks, use your browser's DevTools (Network tab) to inspect the actual requests sent when you click buttons or tabs—replicate those headers and payloads exactly.
内容的提问来源于stack exchange,提问作者pedro regis

