You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python提取维基百科同一行内中印的人口数量及占比

解决维基百科人口表格中印度数据提取失败的问题

问题原因

你的代码通过固定索引data[1]/data[2]/data[3]获取单元格内容,但维基百科的人口列表表格中,部分国家(如印度)的行内单元格包含额外嵌套元素(比如国旗图标、上标注释),导致find_all('td')返回的元素列表索引与预期不符,从而提取到错误数据。

解决方案

改用CSS选择器的nth-child伪类直接定位表格列,不受行内元素结构影响,确保精准获取对应列内容:

import requests
from bs4 import BeautifulSoup

def load_population_dict(url):
    response = requests.get(url)
    soup = BeautifulSoup(response.content, "html.parser")
    
    population_dict = {}
    table = soup.find('table', {'class': 'wikitable'})
    rows = table.find_all('tr')[1:]  # 跳过表头行
    
    for row in rows:
        # 使用nth-child定位对应列,不受行内元素干扰
        country_td = row.select_one('td:nth-child(2)')
        if not country_td:
            continue  # 跳过无效行
        country = country_td.text.strip()
        
        population_td = row.select_one('td:nth-child(3)')
        population = population_td.text.strip() if population_td else 'N/A'
        
        percentage_td = row.select_one('td:nth-child(4)')
        percentage = percentage_td.text.strip() if percentage_td else 'N/A'
        
        population_dict[country] = (population, percentage)
    
    return population_dict

# 测试提取中国和印度数据
pop_dict = load_population_dict("https://en.wikipedia.org/wiki/List_of_countries_and_dependencies_by_population")
print("中国数据:", pop_dict.get("China"))
print("印度数据:", pop_dict.get("India"))

说明

  • td:nth-child(2)对应表格的「国家/地区」列,td:nth-child(3)对应「人口」列,td:nth-child(4)对应「占世界人口比例」列,这是该维基百科表格的固定列位置。
  • 增加空值判断,避免因表格结构异常导致报错。

内容的提问来源于stack exchange,提问作者Kyle

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 14:23:09