如何用Python下载Morningstar网页中的EXPE Key Ratios.csv文件?
Got it, I’ve run into this exact problem before—sites that hide their download links behind JavaScript can feel tricky when you don’t want to mess with parsing JS. Let’s cut straight to the solution for your Morningstar CSV download.
The Short Version: Skip the JS, Grab the Direct Download URL
When you click the "Export" button on that page, your browser sends a simple GET request to a specific endpoint. You don’t need to trace JS code to find it—just use your browser’s dev tools to inspect the network traffic when you hit Export, and you’ll see the real download URL. For your EXPE page, it looks like this:
https://financials.morningstar.com/finan/financials/exportFinancePart.html?&t=EXPE®ion=usa&culture=en-US&cur=&order=asc
Python Code to Automate the Download
Here’s a complete script using requests to fetch and save the CSV file. We’ll add basic headers to mimic a browser request (this helps avoid being blocked by Morningstar’s anti-scraping measures):
import requests # Direct download URL with parameters matching your target page download_url = "https://financials.morningstar.com/finan/financials/exportFinancePart.html?&t=EXPE®ion=usa&culture=en-US&cur=&order=asc" # Mimic a browser's request headers to avoid being blocked headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36", "Referer": "http://financials.morningstar.com/ratios/r.html?t=EXPE®ion=usa&culture=en-US" } # Fetch the CSV content response = requests.get(download_url, headers=headers) response.encoding = "utf-8" # Ensure proper character encoding # Save the content to a local CSV file with open("EXPE Key Ratios.csv", "w", newline="", encoding="utf-8") as csv_file: csv_file.write(response.text) print("Success! 'EXPE Key Ratios.csv' has been downloaded.")
Why This Works
Morningstar’s "Export" button doesn’t generate some super complex hidden URL via JS—it just directs you to that exportFinancePart.html endpoint with the same parameters as the original ratios page (t=EXPE, region=usa, etc.). By replicating that request in Python with the right headers, you bypass the need to interact with the page’s JavaScript entirely.
Notes for Future Use
- If this stops working later, re-check the network tab in your browser dev tools when clicking Export—sometimes sites tweak their endpoint URLs or parameters.
- If you get blocked, try updating the
User-Agentheader to match your current browser’s, or add a small delay between requests if you’re scraping multiple files.
内容的提问来源于stack exchange,提问作者mhsnk

