使用pd.read_html爬取网页表格报错ValueError:未找到表格
First, let's address the immediate issues and root causes:
1. Fix the Filename Syntax Error
In your code, test.csv is missing quotes, which would trigger a NameError once tables are detected. Correct it to:
tables[0].to_csv('test.csv', index=False)
2. Why Pandas Can't Find Tables
Pandas' read_html relies exclusively on standard <table> HTML tags. If the target page uses a non-table layout (like div-based grids) for standings, or returns malformed HTML, pandas won't detect any tables. Additionally, some sites block or alter content for requests without proper browser-like headers.
3. Troubleshooting Steps & Solutions
Step 1: Verify if Tables Exist in the Response
Use BeautifulSoup to check if the page contains <table> elements, and add proper request headers to mimic a browser:
import requests from bs4 import BeautifulSoup url = 'https://www.oaklandyard.com/lg_standings/lg_standings.asp?LgSessCode=2731&ReturnPg=lg%5Fsoccer%5Fcoed%2Easp%232731&ShowRankings=False&HeaderTitle=&sw=1800' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36' } response = requests.get(url, headers=headers) soup = BeautifulSoup(response.text, 'html.parser') # Check for tables in the HTML tables = soup.find_all('table') print(f"Found {len(tables)} tables in the response")
Step 2: If No Tables Are Found (Likely Scenario)
If the output is Found 0 tables, the standings are rendered using non-table elements (e.g., divs with specific classes). You'll need to scrape the data manually:
- Use your browser's dev tools to inspect the standings section and identify container classes/IDs.
- Extract rows and columns from those elements.
Example workflow (adjust based on actual page structure):
# Locate the standings grid (replace with actual class from the page) standings_grid = soup.find('div', class_='standings-container') # Extract all rows rows = standings_grid.find_all('div', class_='standings-row') # Parse data into a list of dictionaries standings_data = [] for row in rows: columns = row.find_all('div', class_='standings-col') team_data = { 'Team Name': columns[0].text.strip(), 'Wins': columns[1].text.strip(), 'Losses': columns[2].text.strip(), 'Ties': columns[3].text.strip(), # Add other columns as needed } standings_data.append(team_data) # Convert to DataFrame and save to CSV import pandas as pd df = pd.DataFrame(standings_data) df.to_csv('oakland_standings.csv', index=False)
Step 3: If Tables Exist But Pandas Can't Parse
If tables are found but read_html fails, extract the table HTML first with BeautifulSoup, then pass it to pandas:
# Get the first table's raw HTML table_html = str(tables[0]) # Parse with pandas df = pd.read_html(table_html)[0] df.to_csv('standings.csv', index=False)
内容的提问来源于stack exchange,提问作者kujoh

