You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Pandas read_html爬维基百科表格,如何修复Population.1列NaN缺失值?

解决维基百科人口表格Population.1列显示NaN的问题

问题原因

你碰到的情况主要是两个问题导致:

  • 直接用requests.get请求时没加浏览器标识,被维基百科判定为爬虫,返回的HTML内容不完整,表格对应列的数据没被正确获取
  • pd.read_html默认的表头解析逻辑和网页表格的实际结构不匹配——该网页的人口表格是复合表头(两行表头),你用header=0只取了第一行表头,导致列对应关系混乱

解决方案

1. 给请求添加浏览器标识,确保获取完整内容

import requests
import pandas as pd

pop_url = "https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population"

# 添加模拟浏览器的请求头,避免被反爬拦截
headers = {
    'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36'
}

r = requests.get(pop_url, headers=headers)

2. 精准定位目标表格并适配复合表头

维基百科的人口表格用了wikitable sortable类,直接指定这个类定位表格避免取错;同时用header=[0,1]解析两行复合表头:

# 定位目标表格并解析复合表头
wiki_tables = pd.read_html(r.text, attrs={'class': 'wikitable sortable'}, header=[0,1])
# 目标表格是第一个(索引0),之前你取的索引1是页面里的其他表格
cont_pop = wiki_tables[0]
cont_pop.head()

3. 可选:用BeautifulSoup手动提取表格(更灵活)

如果pd.read_html还是有解析问题,用BeautifulSoup手动提取表格内容再转成DataFrame:

from bs4 import BeautifulSoup

soup = BeautifulSoup(r.text, 'html.parser')
# 找到目标表格
target_table = soup.find('table', class_='wikitable sortable')

# 逐行提取表格内容
rows = []
for tr in target_table.find_all('tr'):
    row = [td.get_text(strip=True) for td in tr.find_all(['td', 'th'])]
    rows.append(row)

# 转成DataFrame并设置表头(根据实际表头行数调整)
cont_pop = pd.DataFrame(rows[2:], columns=rows[0])
cont_pop.head()

内容的提问来源于stack exchange,提问作者Wisam_Saeed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.19 14:52:27