如何用Python提取维基百科同一行内中印的人口数量及占比
解决维基百科人口表格中印度数据提取失败的问题
问题原因
你的代码通过固定索引data[1]/data[2]/data[3]获取单元格内容,但维基百科的人口列表表格中,部分国家(如印度)的行内单元格包含额外嵌套元素(比如国旗图标、上标注释),导致find_all('td')返回的元素列表索引与预期不符,从而提取到错误数据。
解决方案
改用CSS选择器的nth-child伪类直接定位表格列,不受行内元素结构影响,确保精准获取对应列内容:
import requests from bs4 import BeautifulSoup def load_population_dict(url): response = requests.get(url) soup = BeautifulSoup(response.content, "html.parser") population_dict = {} table = soup.find('table', {'class': 'wikitable'}) rows = table.find_all('tr')[1:] # 跳过表头行 for row in rows: # 使用nth-child定位对应列,不受行内元素干扰 country_td = row.select_one('td:nth-child(2)') if not country_td: continue # 跳过无效行 country = country_td.text.strip() population_td = row.select_one('td:nth-child(3)') population = population_td.text.strip() if population_td else 'N/A' percentage_td = row.select_one('td:nth-child(4)') percentage = percentage_td.text.strip() if percentage_td else 'N/A' population_dict[country] = (population, percentage) return population_dict # 测试提取中国和印度数据 pop_dict = load_population_dict("https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population") print("中国数据:", pop_dict.get("China")) print("印度数据:", pop_dict.get("India"))
说明
td:nth-child(2)对应表格的「国家/地区」列,td:nth-child(3)对应「人口」列,td:nth-child(4)对应「占世界人口比例」列,这是该维基百科表格的固定列位置。- 增加空值判断,避免因表格结构异常导致报错。
内容的提问来源于stack exchange,提问作者Kyle
相关产品推荐
相关产品推荐

