使用Python从URL提取数据时遇KeyError: 'Source'问题求助
问题说明
使用Python的Pandas库从维基百科人口数据页面提取表格,在尝试统计最常见数据源时触发KeyError: 'Source',提示找不到该列。
代码示例
import pandas as pd # 从维基百科国家/地区人口列表页面读取数据到Pandas DataFrame url = "https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population" df = pd.read_html(url, attrs={"class": "wikitable"})[0] # 打印DataFrame内容 print(df) # 打印DataFrame中的记录数 print(f"DataFrame中共有{len(df)}条记录。") # 查找最常见的数据源 most_common_source = df["Source"].value_counts().index[0] print(f"最常见的数据源是{most_common_source}。")
报错信息
KeyError Traceback (most recent call last)
/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3360 try:
-> 3361 return self._engine.get_loc(casted_key) 3362 except KeyError as err:8 frames pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()
pandas/_libs/hashtable_class_helper.pxi in pandas._libs.hashtable.PyObjectHashTable.get_item()
KeyError: 'Source'
The above exception was the direct cause of the following exception:
KeyError Traceback (most recent call last)
/usr/local/lib/python3.8/dist-packages/pandas/core/indexes/base.py in get_loc(self, key, method, tolerance) 3361 return self._engine.get_loc(casted_key) 3362 except KeyError as err:
-> 3363 raise KeyError(key) from err 3364 3365 if is_scalar(key) and isna(key) and not self.hasnans:KeyError: 'Source'
解决步骤
检查实际列名
先打印DataFrame的列名,确认是否存在'Source'列:print(df.columns)维基百科该表格的表头可能是多级结构,或者列名不是英文'Source',实际可能是中文“来源”或其他表述。
扁平化多级表头
若表格存在多级表头,列名会以元组形式存在,需要先扁平化:# 将多级表头合并为单级列名 df.columns = ['_'.join(str(col) for col in col_tuple).strip() for col_tuple in df.columns] # 再次查看列名 print(df.columns)确认目标表格
页面中可能存在多个wikitable类的表格,遍历所有表格找到包含数据源列的那个:tables = pd.read_html(url, attrs={"class": "wikitable"}) for idx, table in enumerate(tables): print(f"表格{idx}的列名:{table.columns}")找到对应表格后,替换
df = tables[0]中的索引值。修正后的完整代码
import pandas as pd url = "https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population" tables = pd.read_html(url, attrs={"class": "wikitable"}) # 选择第一个表格(根据实际情况调整索引) df = tables[0] # 处理多级表头,提取最后一级作为列名 df.columns = [col[-1] if isinstance(col, tuple) else col for col in df.columns] # 确认列名 print(df.columns) # 统计最常见数据源 most_common_source = df["Source"].value_counts().index[0] print(f"最常见的数据源是{most_common_source}。")
内容的提问来源于stack exchange,提问作者Ese Akaunu

