You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Python爬取维基百科计算机科学家信息的技术难题

问题描述

需要编写Python脚本从维基百科计算机科学家列表页面收集每位科学家的以下信息:

  • 全名
  • 获奖数量
  • 就读院校

已完成获取科学家文章链接和提取全名的代码,但无法正确获取获奖数量和就读院校,现有代码片段如下:

获取科学家链接的代码

import requests
from bs4 import BeautifulSoup

URL = "https://en.wikipedia.org/wiki/List_of_computer_scientists"
response = requests.get(URL)
soup = BeautifulSoup(response.content, 'html.parser')
lines = soup.find(id="mw-content-text").find_all("li")

valid_links = []
for line in lines:
    link = line.find("a")
    if link['href'].find("/wiki/") == -1:
        continue
    if link.text == "Lists portal":
        break
    valid_links.append("https://en.wikipedia.org" + link['href'])

信息提取的代码片段

response = requests.get(url)

soup = BeautifulSoup(response.content, 'html.parser')

scientist_name = soup.find(id="firstHeading").string
    
soup.find(id="mw-content-text").find("table", class_="infobox biography vcard")
    
scientist_education = "PLACEHOLDER"
scientist_awards = "PLACEHOLDER"
解决方案

维基百科人物页面的核心信息都在infobox biography vcard表格中,我们可以通过匹配表格内的表头标签(如包含"Education"或"Awards"的行)来提取目标信息,同时兼容不同页面的格式差异。

完整信息提取代码

import requests
from bs4 import BeautifulSoup

def extract_scientist_info(url):
    response = requests.get(url)
    soup = BeautifulSoup(response.content, 'html.parser')
    
    # 提取全名
    scientist_name = soup.find(id="firstHeading").string
    
    # 定位人物信息表格
    infobox = soup.find("table", class_="infobox biography vcard")
    if not infobox:
        return {"name": scientist_name, "education": "无数据", "awards_count": 0}
    
    scientist_education = "无数据"
    awards_count = 0
    
    # 遍历表格行,提取目标信息
    for row in infobox.find_all("tr"):
        header = row.find("th")
        if not header:
            continue
        
        # 提取就读院校:匹配表头含"Education"的行
        if "Education" in header.get_text(strip=True):
            education_cell = row.find("td")
            if education_cell:
                # 提取单元格内所有有效文本(兼容链接和纯文本格式)
                text_items = [item.get_text(strip=True) for item in education_cell.find_all(["a", "span"]) if item.get_text(strip=True)]
                scientist_education = "; ".join(text_items) if text_items else education_cell.get_text(strip=True)
        
        # 统计获奖数量:匹配表头含"Awards"的行
        if "Awards" in header.get_text(strip=True):
            awards_cell = row.find("td")
            if awards_cell:
                # 优先统计列表项数量(多数获奖用<li>展示)
                award_list = awards_cell.find_all("li")
                if award_list:
                    awards_count = len(award_list)
                else:
                    # 无列表时按换行拆分统计
                    award_text = awards_cell.get_text(strip=True)
                    awards_count = len([item for item in award_text.split("\n") if item.strip()]) if award_text else 0
    
    return {
        "name": scientist_name,
        "education": scientist_education,
        "awards_count": awards_count
    }

# 调用示例:遍历已获取的链接提取信息(先测试前5个)
for link in valid_links[:5]:
    info = extract_scientist_info(link)
    print(info)

关键说明

  • 教育信息处理:兼容链接式院校名称和纯文本格式,将多个院校用分号分隔,确保信息完整。
  • 获奖数量统计:优先统计<li>标签数量(最常见的获奖列表格式),无列表时按换行拆分文本统计,适配不同页面的排版差异。
  • 异常兼容:如果页面没有人物信息表格,返回默认的"无数据"和0,避免脚本中断。

内容的提问来源于stack exchange,提问作者George Gkonis

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.02 17:21:06