如何使用Python从指定NSE网页采集公告附件链接并下载附件
Scrape NSE India Corporate Announcements & Download Attachments with Python
Hey Rohit, let's tackle this problem step by step. NSE India's corporate announcements page uses dynamic content, so we have two reliable approaches here: using their official API (the best option for stability) or automating a browser with Selenium. Let's break down both.
Prerequisites
First, install the required libraries via pip:
pip install requests selenium webdriver-manager
Approach 1: Use NSE's Official API (Recommended)
NSE exposes an API endpoint that feeds the corporate announcements data directly. This is far more stable than scraping the UI, which can break if the site's design changes.
Step-by-Step Code
import requests import os from urllib.parse import urljoin # Set up headers to mimic a browser request (NSE blocks plain, unauthenticated requests) headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36', 'Accept': 'application/json, text/plain, */*', 'Referer': 'https://www.nseindia.com/corporates/corporateHome.html' } # API endpoint for equity corporate announcements api_url = 'https://www.nseindia.com/api/corporate-announcements?index=equities' # Create a directory to save attachments (avoids errors if it already exists) save_dir = 'nse_corporate_attachments' os.makedirs(save_dir, exist_ok=True) try: # Fetch data from the API response = requests.get(api_url, headers=headers) response.raise_for_status() # Raise an error for HTTP issues (e.g., 404, 500) announcements = response.json() # Iterate through each announcement to extract attachments for idx, ann in enumerate(announcements.get('data', []), 1): attachment_url = ann.get('attachmentURL') if attachment_url: # Fix relative URLs (some links might not include the full domain) full_url = urljoin('https://www.nseindia.com', attachment_url) # Generate a meaningful filename to avoid overwrites filename = os.path.join(save_dir, f'announcement_{idx}_{os.path.basename(full_url)}') # Download the attachment file_response = requests.get(full_url, headers=headers) file_response.raise_for_status() with open(filename, 'wb') as f: f.write(file_response.content) print(f"Successfully downloaded: {filename}") except Exception as e: print(f"An error occurred during scraping/download: {str(e)}")
Key Notes:
- Headers: NSE blocks requests without proper headers, so we mimic a Chrome browser request to avoid being flagged.
- Relative URLs: Some attachment links are relative, so
urljoinconverts them to full, valid URLs. - Error Handling:
raise_for_status()catches HTTP errors early, so you know if a request fails.
Approach 2: Browser Automation with Selenium
If you need to interact with the UI directly (e.g., apply filters for specific announcements), Selenium simulates a real browser to load dynamic content properly.
Step-by-Step Code
from selenium import webdriver from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from webdriver_manager.chrome import ChromeDriverManager import requests import os # Set up Chrome driver (automatically installs the correct version) driver = webdriver.Chrome(service=Service(ChromeDriverManager().install())) wait = WebDriverWait(driver, 10) # Wait up to 10 seconds for elements to load # Navigate to NSE's corporate home page driver.get('https://www.nseindia.com/corporates/corporateHome.html') try: # Click on the "Corporate Announcements - Equities" tab (wait for it to be clickable) equities_tab = wait.until(EC.element_to_be_clickable((By.LINK_TEXT, 'Corporate Announcements - Equities'))) equities_tab.click() # Wait for the announcements list to load (adjust selector if NSE updates their UI) wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, '.announcement-list'))) # Extract all attachment links from the page attachment_links = driver.find_elements(By.CSS_SELECTOR, '.attachment-link a') attachment_urls = [link.get_attribute('href') for link in attachment_links if link.get_attribute('href')] # Create save directory save_dir = 'nse_attachments_selenium' os.makedirs(save_dir, exist_ok=True) # Download each attachment for idx, url in enumerate(attachment_urls, 1): filename = os.path.join(save_dir, f'announcement_{idx}_{os.path.basename(url)}') response = requests.get(url) with open(filename, 'wb') as f: f.write(response.content) print(f"Downloaded: {filename}") finally: # Always close the browser to avoid hanging processes driver.quit()
Key Notes:
- Dynamic Loading:
WebDriverWaitensures we don't try to interact with elements before they're fully loaded. - Selector Updates: NSE may change their CSS classes over time—if the code breaks, inspect the page to update the selectors.
- Pagination: If there are multiple pages of announcements, add code to click the "Next" button and repeat the extraction loop.
General Best Practices for Web Scraping & File Download
- Respect Rate Limits: Add small delays (
time.sleep(1)) between requests to avoid getting blocked by NSE's servers. - Handle Edge Cases: Add checks for broken links, empty attachments, or unsupported file types.
- Check Robots.txt: Review NSE's
robots.txtto ensure you're allowed to scrape the data you need. - Organize Files: Include metadata like announcement dates or company names in filenames for easier reference.
内容的提问来源于stack exchange,提问作者Rohit
相关产品推荐
相关产品推荐

