如何用Python自动点击Morningstar导出按钮下载财报CSV
First, let's break down why your BeautifulSoup code isn't returning the export button:
Morningstar loads a lot of its page content dynamically using JavaScript. When you use urlopen to fetch the page, you're only getting the initial static HTML source—the export button element you're looking for might not exist in that initial response; it gets injected later by the page's scripts. That's why your soup.find() call comes up empty.
Here are two reliable solutions to automate the CSV export:
Solution 1: Use Selenium to Simulate Manual Browser Actions
This approach mimics exactly what you'd do manually—opening the page, waiting for the button to load, and clicking it. It works perfectly for JavaScript-heavy pages.
Steps & Code:
Install Selenium and your browser's driver (e.g., ChromeDriver for Chrome):
pip install seleniumYou'll also need to download the driver matching your browser version and add it to your system PATH, or specify its path directly in the code.
Write the automation script:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC import time # Optional: Set a custom download directory download_dir = "/path/to/your/download/folder" chrome_options = webdriver.ChromeOptions() prefs = {"download.default_directory": download_dir} chrome_options.add_experimental_option("prefs", prefs) # Initialize the browser driver = webdriver.Chrome(options=chrome_options) try: # Navigate to the target page driver.get("http://financials.morningstar.com/cash-flow/cf.html?t=PIRC®ion=ita&culture=en-US") # Wait up to 10 seconds for the export button to become clickable export_btn = WebDriverWait(driver, 10).until( EC.element_to_be_clickable((By.CLASS_NAME, "rf_export")) ) # Click the export button export_btn.click() # Give the browser time to download the file (adjust based on your internet speed) time.sleep(5) finally: # Close the browser once done driver.quit()
Solution 2: Directly Call the Export API (More Efficient)
Instead of simulating a browser, you can find the actual API endpoint that the export button triggers, then call it directly with requests. This is faster and doesn't require a browser to run.
Steps & Code:
Find the export API URL:
- Open the Morningstar page in your browser, press F12 to open DevTools, go to the Network tab.
- Click the "Export CSV" button, and look for a new request (usually to a URL containing
csvExport.html). - Copy that full request URL, along with any important request headers (like
User-AgentandRefererto avoid being blocked).
Use
requeststo download the CSV:import requests # Replace this with the actual API URL you found in DevTools export_api_url = "http://financials.morningstar.com/export/csvExport.html?t=PIRC®ion=ita&culture=en-US&type=cashflow" # Mimic a browser request to avoid being flagged headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "http://financials.morningstar.com/cash-flow/cf.html?t=PIRC®ion=ita&culture=en-US" } # Send the request and save the CSV response = requests.get(export_api_url, headers=headers) with open("pirc_cash_flow.csv", "wb") as file: file.write(response.content)
Note:
Morningstar might change their API parameters or endpoints over time, so you may need to re-check the network request if this stops working.
内容的提问来源于stack exchange,提问作者M. Az

