使用Python 3实现多日期批量下载印尼证交所股票汇总数据
Automate Bulk Download of IDX Stock Summary Data with Python 3
I’ve worked on similar automation tasks for financial data portals, so here are two reliable approaches to solve your problem—one using browser automation (to mimic the exact user flow) and another using direct API calls (faster and more efficient if feasible).
Method 1: Browser Automation with Selenium
This method replicates the steps you’d manually take (select date, click apply, download data) using Selenium, which is great if you want to stick strictly to the page’s UI workflow.
Prerequisites
- Install Selenium:
pip install selenium - Download the matching webdriver for your browser (e.g., ChromeDriver for Chrome) and ensure it’s in your PATH or specify its path in the code.
Sample Code
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time import os # Configure download directory (replace with your preferred path) DOWNLOAD_DIR = "/path/to/your/download/folder" os.makedirs(DOWNLOAD_DIR, exist_ok=True) # Set up Chrome options to control download behavior chrome_options = webdriver.ChromeOptions() prefs = { "download.default_directory": DOWNLOAD_DIR, "download.prompt_for_download": False, "download.directory_upgrade": True } chrome_options.add_experimental_option("prefs", prefs) # Initialize the driver driver = webdriver.Chrome(options=chrome_options) driver.get("https://www.idx.co.id/en-us/market-data/trading-summary/stock-summary/") # List of dates to download (adjust format to match the page's date picker; e.g., DD/MM/YYYY) TARGET_DATES = ["01/03/2024", "04/03/2024", "05/03/2024"] for date_str in TARGET_DATES: try: # Wait for date picker to load and click it date_picker = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.ID, "datePicker")) # Update selector if needed ) date_picker.click() # Clear existing date and input the target date date_input = WebDriverWait(driver, 10).until( EC.visibility_of_element_located((By.CSS_SELECTOR, "input.form-control.datepicker")) # Adjust selector ) date_input.clear() date_input.send_keys(date_str) # Click the "Apply" button to update the data apply_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Apply')]")) ) apply_btn.click() # Wait for data to load and download button to be clickable download_btn = WebDriverWait(driver, 15).until( EC.element_to_be_clickable((By.XPATH, "//button[contains(text(), 'Download')]")) ) download_btn.click() # Wait for download to complete (adjust sleep time based on your internet speed) time.sleep(6) print(f"Successfully downloaded data for {date_str}") except Exception as e: print(f"Failed to download data for {date_str}: {str(e)}") continue # Clean up: close the browser driver.quit()
Key Notes
- Selectors: Use your browser’s dev tools (Right-click → Inspect) to verify the IDs/XPaths for the date picker, input field, apply button, and download button—these can change if the website updates its UI.
- Date Format: Ensure
date_strmatches the format the date picker expects (check the placeholder text in the input field, e.g.,DD/MM/YYYYorMM/DD/YYYY). - Download Verification: For a more robust solution, replace
time.sleep()with code that checks if the file has finished downloading inDOWNLOAD_DIR.
Method 2: Direct API Call (Faster Alternative)
Most modern websites load data via API calls, which skip the overhead of browser automation. Here’s how to find and use the underlying API:
- Open the IDX stock summary page in your browser.
- Open DevTools (F12) → Go to the Network tab.
- Select a new date and click "Apply"—look for an XHR/fetch request that loads the stock data (it might have a URL containing
stock-summaryormarket-data). - Copy the request URL, headers, and parameters (especially the date parameter, which might be in
YYYY-MM-DDformat).
Sample Code for API Calls
import requests import os DOWNLOAD_DIR = "/path/to/your/download/folder" os.makedirs(DOWNLOAD_DIR, exist_ok=True) # Replace with the actual API endpoint you found (example only) API_URL = "https://api.idx.co.id/market-data/trading-summary/stock-summary" HEADERS = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/122.0.0.0 Safari/537.36", # Add any other required headers (e.g., authorization, referer) from the dev tools request } TARGET_DATES = ["2024-03-01", "2024-03-04", "2024-03-05"] # Use API's expected date format for date_str in TARGET_DATES: try: params = {"date": date_str} response = requests.get(API_URL, headers=HEADERS, params=params) response.raise_for_status() # Raise error for HTTP status codes >=400 # Save the response as a CSV/Excel file (adjust extension based on API response format) file_path = os.path.join(DOWNLOAD_DIR, f"idx_stock_summary_{date_str}.csv") with open(file_path, "wb") as f: f.write(response.content) print(f"Successfully saved data for {date_str} to {file_path}") except Exception as e: print(f"Failed to fetch data for {date_str}: {str(e)}")
Why This Is Better
- Speed: No browser loading times—hundreds of dates can be processed in minutes instead of hours.
- Reliability: Less prone to UI changes (though APIs can also change, they’re often more stable than frontend elements).
Troubleshooting Tips
- Captchas/Blocks: If the website blocks automated requests, try adding delays between requests, rotating user agents, or using
undetected-chromedriverinstead of regular Selenium. - Holiday Dates: The IDX is closed on weekends and public holidays—add logic to skip these dates or handle empty responses gracefully.
- Webdriver Version: Ensure your webdriver version matches your browser version to avoid compatibility errors.
内容的提问来源于stack exchange,提问作者Adhi
相关产品推荐
相关产品推荐

