You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python抓取JS表单:巴西Hemeroteca数据库图片采集技术问询

Hey there! Switching your web scraping tool from C/C++ to Python for team accessibility is a smart move—let’s walk through how to tackle the Hemeroteca Digital database’s cascading forms and grab those image files you need.

1. First, Set Up Your Dependencies

The core tools you’ll need handle HTTP requests, HTML parsing, and dynamic form interactions (since those cascading dropdowns are JavaScript-driven). Install them with:

pip install requests beautifulsoup4 selenium webdriver-manager argparse
2. Core Script: Handle Cascading Forms & Download Images

Since the dropdowns load dynamically (year options only appear after selecting a newspaper, etc.), using Selenium to mimic browser behavior is the most reliable approach. Here’s a commented, reusable script:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import Select, WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import requests
import os
import argparse
import time

def download_images(image_urls, save_dir):
    """批量下载图片到指定文件夹"""
    os.makedirs(save_dir, exist_ok=True)
    for idx, url in enumerate(image_urls):
        try:
            # 加入延迟避免反爬
            time.sleep(1)
            response = requests.get(url, stream=True, timeout=10)
            response.raise_for_status()
            image_path = os.path.join(save_dir, f"issue_{idx+1}.jpg")
            with open(image_path, "wb") as f:
                for chunk in response.iter_content(chunk_size=8192):
                    f.write(chunk)
            print(f"✅ Saved: {image_path}")
        except Exception as e:
            print(f"❌ Failed to download {url}: {str(e)}")

def scrape_hemeroteca(newspaper_name, target_year, save_dir):
    # 初始化浏览器(自动管理Chrome驱动)
    driver = webdriver.Chrome()
    wait = WebDriverWait(driver, 15)

    try:
        # 打开目标页面
        driver.get("http://bndigital.bn.gov.br/hemeroteca-digital/")
        print("📄 Opened Hemeroteca Digital page")

        # 1. 选择报纸/期刊下拉框
        # 注意:替换成页面实际的元素定位(用Chrome F12查看元素ID/XPath)
        newspaper_select = wait.until(
            EC.presence_of_element_located((By.XPATH, "//select[contains(@name, 'jornal')]"))
        )
        select_newspaper = Select(newspaper_select)
        # 按文本选择报纸,也可以用索引或value
        select_newspaper.select_by_visible_text(newspaper_name)
        print(f"📰 Selected newspaper: {newspaper_name}")
        time.sleep(2)  # 给JS时间加载下一个下拉框

        # 2. 选择目标年份
        year_select = wait.until(
            EC.presence_of_element_located((By.XPATH, "//select[contains(@name, 'ano')]"))
        )
        select_year = Select(year_select)
        select_year.select_by_value(str(target_year))
        print(f"📅 Selected year: {target_year}")
        time.sleep(2)

        # 3. 选择期数(这里选第一个可用期数,可改成循环遍历所有期数)
        issue_select = wait.until(
            EC.presence_of_element_located((By.XPATH, "//select[contains(@name, 'numero')]"))
        )
        select_issue = Select(issue_select)
        select_issue.select_by_index(1)  # 跳过第一个默认提示选项
        print("📑 Selected issue")

        # 提交表单(点击搜索按钮)
        submit_btn = wait.until(
            EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Pesquisar')]"))
        )
        submit_btn.click()
        print("🔍 Submitted form, loading issue page...")
        time.sleep(3)

        # 抓取页面中的目标图片(过滤掉无关图片,比如logo)
        image_elements = wait.until(
            EC.presence_of_all_elements_located((By.TAG_NAME, "img"))
        )
        target_image_urls = [
            img.get_attribute("src") for img in image_elements
            if "hemeroteca-digital" in img.get_attribute("src")
        ]
        print(f"🖼️ Found {len(target_image_urls)} images")

        # 下载图片
        download_images(target_image_urls, save_dir)

    finally:
        # 确保浏览器关闭
        driver.quit()
        print("🔚 Browser closed")

if __name__ == "__main__":
    # 命令行参数配置,方便团队成员使用
    parser = argparse.ArgumentParser(description='Scrape images from Hemeroteca Digital')
    parser.add_argument('--newspaper', type=str, required=True, help='Name of the newspaper/jornal (ex: "O Estado de S. Paulo")')
    parser.add_argument('--year', type=int, required=True, help='Target year to scrape')
    parser.add_argument('--save-dir', type=str, default='hemeroteca_images', help='Directory to save images')
    args = parser.parse_args()

    scrape_hemeroteca(args.newspaper, args.year, args.save_dir)
3. Key Tips for Success
  • Element Locators: Use Chrome’s DevTools (F12) to find the actual name, id, or XPath for each dropdown/button—replace the placeholder XPaths in the script with the real ones from the page.
  • Anti-Crawling Measures: The database might block frequent requests, so adding time.sleep() delays (like in the script) helps. For large-scale scraping, consider rotating IPs or using a proxy.
  • Batch Processing: To scrape multiple years/issues, modify the script to loop through year ranges or all options in the issue dropdown.
  • Alternative Static Approach: If you don’t want to use Selenium, inspect the browser’s Network tab (XHR requests) to find the API endpoints that load dropdown options. You can then use requests to call these APIs directly, avoiding browser overhead.
4. Team Sharing Best Practices
  • Requirements File: Create a requirements.txt with exact package versions so your team can install dependencies in one go:
    requests==2.31.0
    beautifulsoup4==4.12.2
    selenium==4.15.2
    webdriver-manager==4.0.1
    argparse==1.4.0
    
  • Documentation: Add comments explaining how to use the script (e.g., python scrape_hemeroteca.py --newspaper "O Estado de S. Paulo" --year 2020).
  • Configurable Options: Extend the script to accept date ranges or multiple newspapers via command line arguments for maximum flexibility.

内容的提问来源于stack exchange,提问作者Phantom139

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 06:50:33