使用Pandas read_html爬维基百科表格,如何修复Population.1列NaN缺失值?
解决维基百科人口表格Population.1列显示NaN的问题
问题原因
你碰到的情况主要是两个问题导致:
- 直接用
requests.get请求时没加浏览器标识,被维基百科判定为爬虫,返回的HTML内容不完整,表格对应列的数据没被正确获取 pd.read_html默认的表头解析逻辑和网页表格的实际结构不匹配——该网页的人口表格是复合表头(两行表头),你用header=0只取了第一行表头,导致列对应关系混乱
解决方案
1. 给请求添加浏览器标识,确保获取完整内容
import requests import pandas as pd pop_url = "https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population" # 添加模拟浏览器的请求头,避免被反爬拦截 headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36' } r = requests.get(pop_url, headers=headers)
2. 精准定位目标表格并适配复合表头
维基百科的人口表格用了wikitable sortable类,直接指定这个类定位表格避免取错;同时用header=[0,1]解析两行复合表头:
# 定位目标表格并解析复合表头 wiki_tables = pd.read_html(r.text, attrs={'class': 'wikitable sortable'}, header=[0,1]) # 目标表格是第一个(索引0),之前你取的索引1是页面里的其他表格 cont_pop = wiki_tables[0] cont_pop.head()
3. 可选:用BeautifulSoup手动提取表格(更灵活)
如果pd.read_html还是有解析问题,用BeautifulSoup手动提取表格内容再转成DataFrame:
from bs4 import BeautifulSoup soup = BeautifulSoup(r.text, 'html.parser') # 找到目标表格 target_table = soup.find('table', class_='wikitable sortable') # 逐行提取表格内容 rows = [] for tr in target_table.find_all('tr'): row = [td.get_text(strip=True) for td in tr.find_all(['td', 'th'])] rows.append(row) # 转成DataFrame并设置表头(根据实际表头行数调整) cont_pop = pd.DataFrame(rows[2:], columns=rows[0]) cont_pop.head()
内容的提问来源于stack exchange,提问作者Wisam_Saeed
相关产品推荐
相关产品推荐

