You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python筛选抓取维基百科表格:提取首位国家元首信息遇IndexError

Fixing Your Wikipedia Heads of State/Government Scraping Issue

Hey there! Let's break down what's going wrong with your code and get that pandas DataFrame you need sorted out.

The Root of the Problem

Your current code has two main issues causing the IndexError: list index out of range and incorrect data extraction:

  1. You're iterating over all <td> elements separately instead of tying them to their corresponding country row. This means you're pulling titles and names from unrelated cells across the entire table.
  2. Not every <td> has two <a> tags—some might have only one, or the structure might differ (like plain text for titles instead of linked text), which triggers the index error when you try to access row.find_all('a')[1].

The Fix: Process Rows as a Whole

We need to loop through each country row, extract the state, then target the first <td> in that row (which contains the head of state info) to pull the title and name together. Here's how to do it properly:

import requests
from bs4 import BeautifulSoup
import pandas as pd

wiki = "https://en.wikipedia.org/wiki/List_of_current_heads_of_state_and_government"
website_url = requests.get(wiki).text
soup = BeautifulSoup(website_url, 'lxml')
my_table = soup.find('table', {'class':'wikitable plainrowheaders'})

states = []
titles = []
names = []

# Iterate through each row (skip the header row with [1:])
for row in my_table.find_all('tr')[1:]:
    # Extract the state name (from the first <a> in the header cell)
    state_cell = row.find('th').find('a')
    if state_cell:
        states.append(state_cell.text.strip())
    else:
        # Handle cases where state name isn't linked (rare, but just in case)
        states.append(row.find('th').text.strip())
    
    # Extract the head of state cell (first <td> in the row)
    head_of_state_cell = row.find('td')
    if head_of_state_cell:
        # Get all linked elements in this cell
        links = head_of_state_cell.find_all('a')
        # Most entries have title first, then name (but let's handle edge cases)
        if len(links) >= 2:
            # Title is first link, name is second
            titles.append(links[0].text.strip())
            names.append(links[1].text.strip())
        elif len(links) == 1:
            # Sometimes title is plain text, name is linked
            # Split the cell text to separate title and name
            cell_text = head_of_state_cell.get_text(separator="|").split("|")
            # Clean up whitespace
            cleaned_text = [text.strip() for text in cell_text if text.strip()]
            if len(cleaned_text) >= 2:
                titles.append(cleaned_text[0])
                names.append(links[0].text.strip())
            else:
                # Fallback if structure is weird
                titles.append("N/A")
                names.append(links[0].text.strip())
        else:
            # No links at all—use plain text
            cell_text = head_of_state_cell.get_text(separator="|").split("|")
            cleaned_text = [text.strip() for text in cell_text if text.strip()]
            if len(cleaned_text) >= 2:
                titles.append(cleaned_text[0])
                names.append(cleaned_text[1])
            else:
                titles.append("N/A")
                names.append("N/A")
    else:
        # Handle empty cells (shouldn't happen in this table)
        titles.append("N/A")
        names.append("N/A")

# Create the DataFrame
df = pd.DataFrame({
    'State': states,
    'Title': titles,
    'Name': names
})

print(df.head())

Key Improvements:

  • Row-level processing: Each state's info is extracted in one loop iteration, so data stays aligned.
  • Error handling: We check the length of <a> tags and handle cases where titles are plain text instead of links, avoiding index errors.
  • Whitespace cleaning: .strip() ensures we don't get extra spaces in our data.
  • Fallback cases: For rare edge cases where the table structure differs, we have defaults to keep the script running.

Notes for Reliability:

Wikipedia tables can change structure over time, so you might want to add more checks or use try-except blocks for extra robustness. For example, some countries might have joint heads of state, but this code handles the most common cases.

内容的提问来源于stack exchange,提问作者aviss

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 10:04:49