You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python数据爬取排障:pd.read_html在for循环中报错No tables found求助

Fixing the ValueError: No tables found in Your Pandas Web Scraping Loop

Hey there! Let's break down why your code works outside the loop but fails inside it, and how to fix it.

The Root Cause

Looking at your code, the key issue lies in how you're constructing the url variable—you're missing the HTTPS protocol prefix (https://) at the start of the URL string.

When you tested pd.read_html(url)[1] outside the loop, you probably used a full, valid URL with https://, so pandas could reach the website successfully. But inside the loop, your generated URLs look like pro-football-reference.com/teams/crd/2020.htm—pandas interprets this as a local file path instead of a web address. Since that local path doesn't exist, it throws the "No tables found" error.

Fixed Code

Here's the corrected version of your code with proper URL formatting and some practical improvements:

import pandas as pd

team_abbrev = ['crd','atl','rav','buf']
year = 2020
list_of_dataframes = []

for team in team_abbrev:
    # Add https:// prefix and use f-strings for cleaner, more readable code
    url = f"https://pro-football-reference.com/teams/{team}/{year}.htm"
    print(url)
    
    # Optional: Add a User-Agent header to avoid being blocked by the website
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}
    
    try:
        df = pd.read_html(url, headers=headers)[1]
        list_of_dataframes.append(df)
        print(f"Successfully scraped data for {team}")
    except IndexError:
        table_count = len(pd.read_html(url, headers=headers))
        print(f"Warning: Only {table_count} table(s) found for {team}—index [1] doesn't exist")
    except ValueError as e:
        print(f"Error scraping {team}: {e}")

Key Improvements Explained

  • Full HTTPS URL: This tells pandas to fetch the webpage from the internet instead of looking for a local file.
  • F-strings: Simplifies URL construction and makes the code easier to maintain.
  • User-Agent Header: Many websites block requests without a valid User-Agent (they flag them as automated bots). Adding this header helps you avoid being blocked.
  • Error Handling: The try-except blocks catch common issues like missing table indexes or unreachable pages, so your loop doesn't crash immediately.

Additional Tips

  • Always print generated URLs to verify they match what you expect before passing them to pd.read_html.
  • If you still run into problems, check if the website uses dynamically loaded tables (which pd.read_html can't parse). In that case, you might need tools like requests + BeautifulSoup or Selenium to scrape the content.

内容的提问来源于stack exchange,提问作者Michael

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 13:07:43