You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用pd.read_html爬取网页表格报错ValueError:未找到表格

Solution: Fixing Table Scraping for Oakland Yard Standings

First, let's address the immediate issues and root causes:

1. Fix the Filename Syntax Error

In your code, test.csv is missing quotes, which would trigger a NameError once tables are detected. Correct it to:

tables[0].to_csv('test.csv', index=False)

2. Why Pandas Can't Find Tables

Pandas' read_html relies exclusively on standard <table> HTML tags. If the target page uses a non-table layout (like div-based grids) for standings, or returns malformed HTML, pandas won't detect any tables. Additionally, some sites block or alter content for requests without proper browser-like headers.

3. Troubleshooting Steps & Solutions

Step 1: Verify if Tables Exist in the Response

Use BeautifulSoup to check if the page contains <table> elements, and add proper request headers to mimic a browser:

import requests
from bs4 import BeautifulSoup

url = 'https://www.oaklandyard.com/lg_standings/lg_standings.asp?LgSessCode=2731&ReturnPg=lg%5Fsoccer%5Fcoed%2Easp%232731&ShowRankings=False&HeaderTitle=&sw=1800'
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'
}
response = requests.get(url, headers=headers)
soup = BeautifulSoup(response.text, 'html.parser')

# Check for tables in the HTML
tables = soup.find_all('table')
print(f"Found {len(tables)} tables in the response")

Step 2: If No Tables Are Found (Likely Scenario)

If the output is Found 0 tables, the standings are rendered using non-table elements (e.g., divs with specific classes). You'll need to scrape the data manually:

  1. Use your browser's dev tools to inspect the standings section and identify container classes/IDs.
  2. Extract rows and columns from those elements.

Example workflow (adjust based on actual page structure):

# Locate the standings grid (replace with actual class from the page)
standings_grid = soup.find('div', class_='standings-container')

# Extract all rows
rows = standings_grid.find_all('div', class_='standings-row')

# Parse data into a list of dictionaries
standings_data = []
for row in rows:
    columns = row.find_all('div', class_='standings-col')
    team_data = {
        'Team Name': columns[0].text.strip(),
        'Wins': columns[1].text.strip(),
        'Losses': columns[2].text.strip(),
        'Ties': columns[3].text.strip(),
        # Add other columns as needed
    }
    standings_data.append(team_data)

# Convert to DataFrame and save to CSV
import pandas as pd
df = pd.DataFrame(standings_data)
df.to_csv('oakland_standings.csv', index=False)

Step 3: If Tables Exist But Pandas Can't Parse

If tables are found but read_html fails, extract the table HTML first with BeautifulSoup, then pass it to pandas:

# Get the first table's raw HTML
table_html = str(tables[0])
# Parse with pandas
df = pd.read_html(table_html)[0]
df.to_csv('standings.csv', index=False)

内容的提问来源于stack exchange,提问作者kujoh

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 16:37:30