You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python3抓取BNMP网站带搜索筛选的多页动态表格数据?

Scraping Dynamic BNMP Table Data for Rio de Janeiro with Python

Hey there! To scrape the 53k+ entries from the BNMP site's dynamic table, we'll use Selenium (to handle JavaScript-rendered content) and Pandas (to store the data in a DataFrame). Here's a step-by-step solution:

Prerequisites

First, install the required packages:

pip install selenium pandas webdriver-manager

Full Code Implementation

This script will automate selecting Rio de Janeiro, clicking "Pesquisar", paginating through results, and collecting all data:

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from webdriver_manager.chrome import ChromeDriverManager
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import pandas as pd
import time
import random

# Initialize Chrome driver (automatically installs ChromeDriver)
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
driver.maximize_window()
driver.get("http://www.cnj.jus.br/bnmp/#/pesquisar")

try:
    # Wait for the Estado dropdown to load and open it (handles custom React-select component)
    estado_dropdown = WebDriverWait(driver, 15).until(
        EC.element_to_be_clickable((By.XPATH, "//div[contains(text(), 'Estado')]/following-sibling::div//div[contains(@class, 'select__control')]"))
    )
    estado_dropdown.click()

    # Select Rio de Janeiro from the dropdown options
    rio_option = WebDriverWait(driver, 15).until(
        EC.element_to_be_clickable((By.XPATH, "//div[contains(@class, 'select__option') and text()='Rio de Janeiro']"))
    )
    rio_option.click()

    # Click the Pesquisar button to load results
    pesquisar_btn = WebDriverWait(driver, 15).until(
        EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Pesquisar')]"))
    )
    pesquisar_btn.click()

    # Wait for the results table to appear
    WebDriverWait(driver, 20).until(
        EC.presence_of_element_located((By.XPATH, "//table[contains(@class, 'table')]/tbody"))
    )

    # Initialize list to store all data entries
    all_data = []
    page_count = 0

    while True:
        page_count += 1
        print(f"Scraping page {page_count}...")

        # Wait for rows to load and extract all visible rows
        rows = WebDriverWait(driver, 15).until(
            EC.presence_of_all_elements_located((By.XPATH, "//table[contains(@class, 'table')]/tbody/tr"))
        )

        # Extract data from each row
        for row in rows:
            try:
                numero = row.find_element(By.XPATH, "./td[1]").text.strip()
                nome = row.find_element(By.XPATH, "./td[2]").text.strip()
                situacao = row.find_element(By.XPATH, "./td[3]").text.strip()

                # Only add non-empty entries (skip any empty rows)
                if numero and nome:
                    all_data.append({
                        "Numero": numero,
                        "Nome": nome,
                        "Situacao": situacao
                    })
            except Exception as e:
                print(f"Error extracting row: {e}")
                continue

        # Check if next page button exists and is clickable
        try:
            next_btn = WebDriverWait(driver, 5).until(
                EC.element_to_be_clickable((By.XPATH, "//li[@class='page-item next']/a"))
            )
            # Check if the button is disabled (no more pages)
            if "disabled" in next_btn.get_attribute("class"):
                break
            # Click next page with random delay to avoid rate limiting
            next_btn.click()
            time.sleep(random.uniform(1.5, 3.0))
        except:
            # No more pages to scrape
            break

    # Convert collected data to DataFrame
    df = pd.DataFrame(all_data)
    print(f"\nScraping complete! Collected {len(df)} entries.")
    print(df.head())

    # Optional: Save to CSV
    df.to_csv("bnmp_rio_de_janeiro_data.csv", index=False, encoding="utf-8")
    print("Data saved to bnmp_rio_de_janeiro_data.csv")

finally:
    # Close the browser
    driver.quit()

Key Notes & Troubleshooting

  • Dynamic Content Handling: Selenium is used because the table loads data via JavaScript—simple requests won't work here.
  • XPath Selectors: The selectors are tailored to the site's current HTML structure. If the site updates, you may need to adjust them using browser dev tools (F12) to inspect elements.
  • Rate Limiting: Random delays between page clicks help avoid triggering anti-scraping measures. If you get blocked, increase the delay or add a longer pause between pages.
  • Error Handling: The script includes basic error handling to skip problematic rows and continue scraping.
  • Performance: Scraping 53k entries will take time (expect 10-30 minutes depending on your internet speed). You can monitor progress via the console prints.

内容的提问来源于stack exchange,提问作者Reinaldo Chaves

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 10:10:15