已获取网页中Excel文件href链接,如何用Python打开对应工作簿?
Alright, let's build on the code you already have to grab that Excel file and open it up. Here's a step-by-step breakdown with working code that fits your needs:
Step 1: Filter Out the Correct Excel Link
First, instead of printing every link, we'll narrow down to the one that ends with .xls (our target file type). We'll also fix the relative href to a full valid URL since the links from the page are likely not complete.
Step 2: Download the Excel File
We'll use requests to fetch the file content—you can either save it to your local machine or process it directly in memory (no need to save a file if you don't want to).
Step 3: Open and Interact with the Workbook
We'll use openpyxl (a widely used library for Excel handling) to load the workbook. If you don't have it installed yet, run pip install openpyxl in your terminal first.
Full Working Code
import requests from bs4 import BeautifulSoup from openpyxl import load_workbook from io import BytesIO # Fetch the target page page = requests.get("https://fred.stlouisfed.org/series/IC4WSA") soup = BeautifulSoup(page.content, 'html.parser') # Locate the Excel (.xls) link excel_href = None for link in soup.find_all('a'): href = link.get('href') # Check if the link points to an Excel file if href and href.endswith('.xls'): excel_href = href break # Stop searching once we find the first valid Excel link if not excel_href: print("No Excel file link found on the page.") else: # Build the full URL from the relative href full_excel_url = f"https://fred.stlouisfed.org{excel_href}" # Download the Excel file content excel_response = requests.get(full_excel_url) excel_response.raise_for_status() # Catch any download errors (like 404) # Option 1: Save the file to your local disk with open("ic4wsa_data.xls", "wb") as f: f.write(excel_response.content) print("Excel file saved locally as 'ic4wsa_data.xls'") # Option 2: Open the workbook directly from memory (no local file needed) workbook = load_workbook(filename=BytesIO(excel_response.content), read_only=True) # Quick check to confirm the workbook loaded print("Loaded workbook sheet names:", workbook.sheetnames) # Example: Access the first sheet and read a cell value active_sheet = workbook.active print("Value of cell A1:", active_sheet['A1'].value)
Bonus: Use Pandas for Data Analysis
If you're working with the data inside the Excel file and want to analyze it, you can replace the openpyxl section with pandas (install via pip install pandas openpyxl):
import pandas as pd df = pd.read_excel(BytesIO(excel_response.content)) print("First 5 rows of data:\n", df.head())
内容的提问来源于stack exchange,提问作者SBF12345

