You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python从指定NSE网页采集公告附件链接并下载附件

Scrape NSE India Corporate Announcements & Download Attachments with Python

Hey Rohit, let's tackle this problem step by step. NSE India's corporate announcements page uses dynamic content, so we have two reliable approaches here: using their official API (the best option for stability) or automating a browser with Selenium. Let's break down both.

Prerequisites

First, install the required libraries via pip:

pip install requests selenium webdriver-manager

NSE exposes an API endpoint that feeds the corporate announcements data directly. This is far more stable than scraping the UI, which can break if the site's design changes.

Step-by-Step Code

import requests
import os
from urllib.parse import urljoin

# Set up headers to mimic a browser request (NSE blocks plain, unauthenticated requests)
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36',
    'Accept': 'application/json, text/plain, */*',
    'Referer': 'https://www.nseindia.com/corporates/corporateHome.html'
}

# API endpoint for equity corporate announcements
api_url = 'https://www.nseindia.com/api/corporate-announcements?index=equities'

# Create a directory to save attachments (avoids errors if it already exists)
save_dir = 'nse_corporate_attachments'
os.makedirs(save_dir, exist_ok=True)

try:
    # Fetch data from the API
    response = requests.get(api_url, headers=headers)
    response.raise_for_status()  # Raise an error for HTTP issues (e.g., 404, 500)
    announcements = response.json()

    # Iterate through each announcement to extract attachments
    for idx, ann in enumerate(announcements.get('data', []), 1):
        attachment_url = ann.get('attachmentURL')
        if attachment_url:
            # Fix relative URLs (some links might not include the full domain)
            full_url = urljoin('https://www.nseindia.com', attachment_url)
            # Generate a meaningful filename to avoid overwrites
            filename = os.path.join(save_dir, f'announcement_{idx}_{os.path.basename(full_url)}')
            
            # Download the attachment
            file_response = requests.get(full_url, headers=headers)
            file_response.raise_for_status()
            
            with open(filename, 'wb') as f:
                f.write(file_response.content)
            print(f"Successfully downloaded: {filename}")

except Exception as e:
    print(f"An error occurred during scraping/download: {str(e)}")

Key Notes:

  • Headers: NSE blocks requests without proper headers, so we mimic a Chrome browser request to avoid being flagged.
  • Relative URLs: Some attachment links are relative, so urljoin converts them to full, valid URLs.
  • Error Handling: raise_for_status() catches HTTP errors early, so you know if a request fails.

Approach 2: Browser Automation with Selenium

If you need to interact with the UI directly (e.g., apply filters for specific announcements), Selenium simulates a real browser to load dynamic content properly.

Step-by-Step Code

from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from webdriver_manager.chrome import ChromeDriverManager
import requests
import os

# Set up Chrome driver (automatically installs the correct version)
driver = webdriver.Chrome(service=Service(ChromeDriverManager().install()))
wait = WebDriverWait(driver, 10)  # Wait up to 10 seconds for elements to load

# Navigate to NSE's corporate home page
driver.get('https://www.nseindia.com/corporates/corporateHome.html')

try:
    # Click on the "Corporate Announcements - Equities" tab (wait for it to be clickable)
    equities_tab = wait.until(EC.element_to_be_clickable((By.LINK_TEXT, 'Corporate Announcements - Equities')))
    equities_tab.click()

    # Wait for the announcements list to load (adjust selector if NSE updates their UI)
    wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '.announcement-list')))

    # Extract all attachment links from the page
    attachment_links = driver.find_elements(By.CSS_SELECTOR, '.attachment-link a')
    attachment_urls = [link.get_attribute('href') for link in attachment_links if link.get_attribute('href')]

    # Create save directory
    save_dir = 'nse_attachments_selenium'
    os.makedirs(save_dir, exist_ok=True)

    # Download each attachment
    for idx, url in enumerate(attachment_urls, 1):
        filename = os.path.join(save_dir, f'announcement_{idx}_{os.path.basename(url)}')
        response = requests.get(url)
        with open(filename, 'wb') as f:
            f.write(response.content)
        print(f"Downloaded: {filename}")

finally:
    # Always close the browser to avoid hanging processes
    driver.quit()

Key Notes:

  • Dynamic Loading: WebDriverWait ensures we don't try to interact with elements before they're fully loaded.
  • Selector Updates: NSE may change their CSS classes over time—if the code breaks, inspect the page to update the selectors.
  • Pagination: If there are multiple pages of announcements, add code to click the "Next" button and repeat the extraction loop.

General Best Practices for Web Scraping & File Download

  • Respect Rate Limits: Add small delays (time.sleep(1)) between requests to avoid getting blocked by NSE's servers.
  • Handle Edge Cases: Add checks for broken links, empty attachments, or unsupported file types.
  • Check Robots.txt: Review NSE's robots.txt to ensure you're allowed to scrape the data you need.
  • Organize Files: Include metadata like announcement dates or company names in filenames for easier reference.

内容的提问来源于stack exchange,提问作者Rohit

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.11 09:27:17