You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何抓取带有多表头的表格?解决PyCharm中表格抓取出现KeyError的问题

Fixing KeyError and Handling Multi-Header Tables in Web Scraping with Pandas

Let's tackle your two main issues: resolving the KeyError in your current code and learning how to scrape tables with multi-level headers.

First: Fixing the KeyError in Your Code

The KeyError is almost certainly happening because one or more column names in your display_data list don't exist in the DataFrame after renaming, or the columns you're trying to rename (Unnamed: 2, Unnamed: 3) aren't present in the table you're targeting (url[1]).

Here's how to debug and fix it:

  1. First, inspect the actual columns of the table you're loading
    Before doing any renaming, print out the columns of url[1] to see what you're working with:

    import pandas as pd
    url = pd.read_html("https://azurlane.koumakan.jp/List_of_Large_Cruiser_Guns")
    df = url[1]
    print(df.columns) # This will show you the real column names
    

    This will tell you if Unnamed: 2 or Unnamed: 3 actually exist, and what other columns are available.

  2. Adjust your renaming and display columns to match reality
    Let's assume after inspecting, the columns are as expected, but maybe some of your display_data columns are misnamed. For example, maybe the column is called Rounds instead of Rnd, or Damage instead of Dmg. Update the table_headers and display_data lists to match the actual column names from the printout.

  3. A revised version of your code (with safeguards)
    Here's a modified version that adds checks to avoid KeyError:

    import pandas as pd
    pd.set_option('display.max_columns', None)
    pd.set_option('display.width', None)
    pd.set_option('display.max_colwidth', None)
    
    # Load tables and pick the right one
    tables = pd.read_html("https://azurlane.koumakan.jp/List_of_Large_Cruiser_Guns")
    df = tables[1]
    
    # First, print columns to verify
    print("Original columns:", df.columns.tolist())
    
    # Rename columns - adjust these to match the actual unnamed columns you see
    table_headers = {'Unnamed: 2': 'Firepower', 'Unnamed: 3': 'Anti-air'}
    df.rename(columns=table_headers, inplace=True)
    
    # Fill missing values and reset index
    df.fillna('0', inplace=True)
    df.reset_index(drop=True, inplace=True)
    
    # User input
    eq_name = input("Enter equipment name: ").casefold()
    
    # Filter the row
    if eq_name == "triple 283mm (sk c/28)":
        df = df.loc[[0]]
    
    # Ensure only existing columns are selected
    available_columns = [col for col in ['Equipment', 'Firepower', 'Anti-air', 'Rnd', 'Dmg', 'Coef', 'VT', 'Rld', 'Surface DPS', 'Rng', 'Sprd', 'Angle', 'Attr', 'Ammo'] if col in df.columns]
    name_notes = df[available_columns]
    print(name_notes)
    

    The key change here is the available_columns list, which filters out any column names that don't exist in the DataFrame, preventing KeyError.

Second: Scraping Tables with Multi-Level Headers

Many wiki tables (like the Azur Lane one) use multi-row headers. Here's how to handle them properly with pandas:

  1. Use the header parameter in pd.read_html
    If the table has headers spanning multiple rows, specify which rows to use as headers. For example, if the header uses the first 2 rows:

    tables = pd.read_html("your_url_here", header=[0,1])
    

    This will create a MultiIndex for the columns, where each column has a tuple of header values.

  2. Flatten multi-level headers
    If you want to combine the multi-level headers into a single string for easier access, you can flatten them:

    df = tables[0]
    # Combine multi-level headers into one string
    df.columns = ['_'.join(col).strip() for col in df.columns.values]
    

    For example, if a column has headers ['Stats', 'Firepower'], it becomes 'Stats_Firepower'.

  3. Example for multi-header tables
    Let's say you're scraping a table with 2 header rows. Here's a complete snippet:

    import pandas as pd
    
    # Load table with multi-row headers
    tables = pd.read_html("https://example.com/multi-header-table", header=[0,1])
    df = tables[0]
    
    # Flatten the headers
    df.columns = [' '.join(filter(None, col)).strip() for col in df.columns]
    
    # Now you can access columns with their flattened names
    print(df.columns)
    

Final Notes

  • Always inspect the raw columns of the table you're scraping first—this will save you from guessing column names and hitting KeyErrors.
  • For wiki tables, the structure can change, so adding checks (like the available_columns filter) will make your code more robust.

内容的提问来源于stack exchange,提问作者Coelll

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.29 16:24:07