使用Python从网站下载.ods文件失败(仅获取HTML文档)的解决方案求助
I see the core issue here — you’re trying to download content from an Office Online preview page instead of the direct URL for the .ods file itself. Let’s break down why this fails and how to fix it.
Why Your Current Code Isn’t Working
The link you’re using (https://view.officeapps.live.com/op/view.aspx?src=...) is a wrapper for Office’s online preview tool, not a direct file download link. When you send requests to this URL, you’re getting the HTML for the preview interface (the "We're fetching your file..." page you saw) instead of the actual .ods content. This is why read_ods throws a KeyError: '.ods' — it’s trying to parse HTML as an ODS file, which isn’t valid.
The Fix: Use the Direct File URL
The real .ods file address is hidden in the src query parameter of the preview link. If you decode that parameter, you’ll get:https://www.statistik.at/fileadmin/pages/214/2_Verbraucherpreisindizes_ab_1990.ods
Directly requesting this URL will give you the actual file, not the preview HTML.
Corrected Code Example
import requests from pandas_ods_reader import read_ods # Direct URL to the .ods file file_url = "https://www.statistik.at/fileadmin/pages/214/2_Verbraucherpreisindizes_ab_1990.ods" # Optional: Add a User-Agent header to mimic a browser (some servers block non-browser requests) headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } # Download the file response = requests.get(file_url, headers=headers) response.raise_for_status() # Throw an error if the request fails (e.g., 404, 500) # Save the file locally with open("2_Verbraucherpreisindizes_ab_1990.ods", "wb") as out_file: out_file.write(response.content) # Read the ODS file into a DataFrame df = read_ods("2_Verbraucherpreisindizes_ab_1990.ods", sheet=1) # sheet=1 targets the second sheet (0-indexed) print(df.head())
Bonus: Automatically Extract the File URL (For Dynamic Links)
If you need to handle similar Office Online preview links in the future, you can programmatically extract the src parameter:
from urllib.parse import urlparse, parse_qs preview_url = "https://view.officeapps.live.com/op/view.aspx?src=https%3A%2F%2Fwww.statistik.at%2Ffileadmin%2Fpages%2F214%2F2_Verbraucherpreisindizes_ab_1990.ods" # Parse the preview URL to get query parameters parsed_url = urlparse(preview_url) query_params = parse_qs(parsed_url.query) actual_file_url = query_params["src"][0] # Extract the decoded file URL # Use actual_file_url for downloading as shown above
Key Notes
- Always check the response status code (using
response.raise_for_status()) to catch issues like broken links or access restrictions. - Ensure you have the latest versions of required libraries: run
pip install --upgrade requests pandas-ods-readerto update. - The User-Agent header helps avoid being blocked by servers that restrict non-browser requests.
内容的提问来源于stack exchange,提问作者Paul Strasser

