如何在Python中通过HTML导入网页表格?含具体实操疑问
Hey Tom, nice question! Let's walk through two straightforward ways to turn that embedded table into a CSV file—one using pandas (super easy) and another with BeautifulSoup if you prefer more control.
Method 1: Use Pandas (Quickest Way)
Pandas has a built-in read_html() function that can automatically parse HTML tables into DataFrames, which we can then save directly as CSV. Here's how:
import pandas as pd # Replace this with your actual embedded table URL embed_url = "YOUR_EMBED_TABLE_URL" # Fetch and parse the table(s) from the URL # Sports Reference's embedded tables usually show up as the first item in the list df = pd.read_html(embed_url)[0] # Clean up duplicate header rows (Sports Reference repeats headers when scrolling) df = df.drop_duplicates(subset=df.columns[0], keep="first") # Save to CSV df.to_csv("landry_fields_advanced_stats.csv", index=False)
Notes:
- If you run into access issues (like a 403 error), add a user-agent header to mimic a browser:
headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} df = pd.read_html(embed_url, headers=headers)[0] - The
drop_duplicatesstep removes those repeated header rows that are common on Sports Reference tables.
Method 2: Use BeautifulSoup + CSV Module (More Control)
If you want to handle the parsing manually, you can use BeautifulSoup to extract table data and write it to CSV directly:
import requests from bs4 import BeautifulSoup import csv embed_url = "YOUR_EMBED_TABLE_URL" # Fetch the page content with a user-agent to avoid blocks headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"} response = requests.get(embed_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") # Locate the target table table = soup.find("table") # Write to CSV with open("landry_fields_advanced_stats.csv", "w", newline="", encoding="utf-8") as csv_file: writer = csv.writer(csv_file) for row in table.find_all("tr"): # Extract text from each cell, stripping extra whitespace cells = [cell.get_text(strip=True) for cell in row.find_all(["th", "td"])] # Skip empty rows if cells: writer.writerow(cells)
Bonus Tip:
Did you know Sports Reference often has hidden CSV export options? If you go back to the original player page, look for the "Share & More" button (you already used this for embedding)—sometimes there's an "Export" option that lets you download the CSV directly without writing code! But if that's not available for the specific table you want, the two methods above work perfectly.
内容的提问来源于stack exchange,提问作者Tom

