You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python从URL提取数据时遇KeyError: 'Source'问题求助

Pandas读取维基百科表格时KeyError: 'Source'问题解决

问题说明

使用Python的Pandas库从维基百科人口数据页面提取表格,在尝试统计最常见数据源时触发KeyError: 'Source',提示找不到该列。

代码示例

import pandas as pd

# 从维基百科国家/地区人口列表页面读取数据到Pandas DataFrame
url = "https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population"
df = pd.read_html(url, attrs={"class": "wikitable"})[0]

# 打印DataFrame内容
print(df)

# 打印DataFrame中的记录数
print(f"DataFrame中共有{len(df)}条记录。")

# 查找最常见的数据源
most_common_source = df["Source"].value_counts().index[0]

print(f"最常见的数据源是{most_common_source}。")

报错信息

KeyError Traceback (most recent call last)
/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3360 try:
-> 3361 return self._engine.get_loc(casted_key) 3362 except KeyError as err:

8 frames pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()

KeyError: 'Source'

The above exception was the direct cause of the following exception:

KeyError Traceback (most recent call last)
/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3361 return self._engine.get_loc(casted_key) 3362 except KeyError as err:
-> 3363 raise KeyError(key) from err 3364 3365 if is_scalar(key) and isna(key) and not self.hasnans:

KeyError: 'Source'

解决步骤

  1. 检查实际列名
    先打印DataFrame的列名,确认是否存在'Source'列:

    print(df.columns)
    

    维基百科该表格的表头可能是多级结构,或者列名不是英文'Source',实际可能是中文“来源”或其他表述。

  2. 扁平化多级表头
    若表格存在多级表头,列名会以元组形式存在,需要先扁平化:

    # 将多级表头合并为单级列名
    df.columns = ['_'.join(str(col) for col in col_tuple).strip() for col_tuple in df.columns]
    # 再次查看列名
    print(df.columns)
    
  3. 确认目标表格
    页面中可能存在多个wikitable类的表格,遍历所有表格找到包含数据源列的那个:

    tables = pd.read_html(url, attrs={"class": "wikitable"})
    for idx, table in enumerate(tables):
        print(f"表格{idx}的列名:{table.columns}")
    

    找到对应表格后,替换df = tables[0]中的索引值。

  4. 修正后的完整代码

    import pandas as pd
    
    url = "https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population"
    tables = pd.read_html(url, attrs={"class": "wikitable"})
    # 选择第一个表格(根据实际情况调整索引)
    df = tables[0]
    # 处理多级表头,提取最后一级作为列名
    df.columns = [col[-1] if isinstance(col, tuple) else col for col in df.columns]
    # 确认列名
    print(df.columns)
    # 统计最常见数据源
    most_common_source = df["Source"].value_counts().index[0]
    print(f"最常见的数据源是{most_common_source}。")
    

内容的提问来源于stack exchange,提问作者Ese Akaunu

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 20:15:43