如何用Python3抓取BNMP网站带搜索筛选的多页动态表格数据?
Scraping Dynamic BNMP Table Data for Rio de Janeiro with Python
Hey there! To scrape the 53k+ entries from the BNMP site's dynamic table, we'll use Selenium (to handle JavaScript-rendered content) and Pandas (to store the data in a DataFrame). Here's a step-by-step solution:
Prerequisites
First, install the required packages:
pip install selenium pandas webdriver-manager
Full Code Implementation
This script will automate selecting Rio de Janeiro, clicking "Pesquisar", paginating through results, and collecting all data:
from selenium import webdriver from selenium.webdriver.chrome.service import Service from webdriver_manager.chrome import ChromeDriverManager from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import pandas as pd import time import random # Initialize Chrome driver (automatically installs ChromeDriver) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) driver.maximize_window() driver.get("http://www.cnj.jus.br/bnmp/#/pesquisar") try: # Wait for the Estado dropdown to load and open it (handles custom React-select component) estado_dropdown = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.XPATH, "//div[contains(text(), 'Estado')]/following-sibling::div//div[contains(@class, 'select__control')]")) ) estado_dropdown.click() # Select Rio de Janeiro from the dropdown options rio_option = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'select__option') and text()='Rio de Janeiro']")) ) rio_option.click() # Click the Pesquisar button to load results pesquisar_btn = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Pesquisar')]")) ) pesquisar_btn.click() # Wait for the results table to appear WebDriverWait(driver, 20).until( EC.presence_of_element_located((By.XPATH, "//table[contains(@class, 'table')]/tbody")) ) # Initialize list to store all data entries all_data = [] page_count = 0 while True: page_count += 1 print(f"Scraping page {page_count}...") # Wait for rows to load and extract all visible rows rows = WebDriverWait(driver, 15).until( EC.presence_of_all_elements_located((By.XPATH, "//table[contains(@class, 'table')]/tbody/tr")) ) # Extract data from each row for row in rows: try: numero = row.find_element(By.XPATH, "./td[1]").text.strip() nome = row.find_element(By.XPATH, "./td[2]").text.strip() situacao = row.find_element(By.XPATH, "./td[3]").text.strip() # Only add non-empty entries (skip any empty rows) if numero and nome: all_data.append({ "Numero": numero, "Nome": nome, "Situacao": situacao }) except Exception as e: print(f"Error extracting row: {e}") continue # Check if next page button exists and is clickable try: next_btn = WebDriverWait(driver, 5).until( EC.element_to_be_clickable((By.XPATH, "//li[@class='page-item next']/a")) ) # Check if the button is disabled (no more pages) if "disabled" in next_btn.get_attribute("class"): break # Click next page with random delay to avoid rate limiting next_btn.click() time.sleep(random.uniform(1.5, 3.0)) except: # No more pages to scrape break # Convert collected data to DataFrame df = pd.DataFrame(all_data) print(f"\nScraping complete! Collected {len(df)} entries.") print(df.head()) # Optional: Save to CSV df.to_csv("bnmp_rio_de_janeiro_data.csv", index=False, encoding="utf-8") print("Data saved to bnmp_rio_de_janeiro_data.csv") finally: # Close the browser driver.quit()
Key Notes & Troubleshooting
- Dynamic Content Handling: Selenium is used because the table loads data via JavaScript—simple
requestswon't work here. - XPath Selectors: The selectors are tailored to the site's current HTML structure. If the site updates, you may need to adjust them using browser dev tools (F12) to inspect elements.
- Rate Limiting: Random delays between page clicks help avoid triggering anti-scraping measures. If you get blocked, increase the delay or add a longer pause between pages.
- Error Handling: The script includes basic error handling to skip problematic rows and continue scraping.
- Performance: Scraping 53k entries will take time (expect 10-30 minutes depending on your internet speed). You can monitor progress via the console prints.
内容的提问来源于stack exchange,提问作者Reinaldo Chaves
相关产品推荐
相关产品推荐

