使用Python爬取维基百科计算机科学家信息的技术难题
问题描述
需要编写Python脚本从维基百科计算机科学家列表页面收集每位科学家的以下信息:
- 全名
- 获奖数量
- 就读院校
已完成获取科学家文章链接和提取全名的代码,但无法正确获取获奖数量和就读院校,现有代码片段如下:
获取科学家链接的代码
import requests from bs4 import BeautifulSoup URL = "https://en.wikipedia.org/wiki/List_of_computer_scientists" response = requests.get(URL) soup = BeautifulSoup(response.content, 'html.parser') lines = soup.find(id="mw-content-text").find_all("li") valid_links = [] for line in lines: link = line.find("a") if link['href'].find("/wiki/") == -1: continue if link.text == "Lists portal": break valid_links.append("https://en.wikipedia.org" + link['href'])
信息提取的代码片段
response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') scientist_name = soup.find(id="firstHeading").string soup.find(id="mw-content-text").find("table", class_="infobox biography vcard") scientist_education = "PLACEHOLDER" scientist_awards = "PLACEHOLDER"
解决方案
维基百科人物页面的核心信息都在infobox biography vcard表格中,我们可以通过匹配表格内的表头标签(如包含"Education"或"Awards"的行)来提取目标信息,同时兼容不同页面的格式差异。
完整信息提取代码
import requests from bs4 import BeautifulSoup def extract_scientist_info(url): response = requests.get(url) soup = BeautifulSoup(response.content, 'html.parser') # 提取全名 scientist_name = soup.find(id="firstHeading").string # 定位人物信息表格 infobox = soup.find("table", class_="infobox biography vcard") if not infobox: return {"name": scientist_name, "education": "无数据", "awards_count": 0} scientist_education = "无数据" awards_count = 0 # 遍历表格行,提取目标信息 for row in infobox.find_all("tr"): header = row.find("th") if not header: continue # 提取就读院校:匹配表头含"Education"的行 if "Education" in header.get_text(strip=True): education_cell = row.find("td") if education_cell: # 提取单元格内所有有效文本(兼容链接和纯文本格式) text_items = [item.get_text(strip=True) for item in education_cell.find_all(["a", "span"]) if item.get_text(strip=True)] scientist_education = "; ".join(text_items) if text_items else education_cell.get_text(strip=True) # 统计获奖数量:匹配表头含"Awards"的行 if "Awards" in header.get_text(strip=True): awards_cell = row.find("td") if awards_cell: # 优先统计列表项数量(多数获奖用<li>展示) award_list = awards_cell.find_all("li") if award_list: awards_count = len(award_list) else: # 无列表时按换行拆分统计 award_text = awards_cell.get_text(strip=True) awards_count = len([item for item in award_text.split("\n") if item.strip()]) if award_text else 0 return { "name": scientist_name, "education": scientist_education, "awards_count": awards_count } # 调用示例:遍历已获取的链接提取信息(先测试前5个) for link in valid_links[:5]: info = extract_scientist_info(link) print(info)
关键说明
- 教育信息处理:兼容链接式院校名称和纯文本格式,将多个院校用分号分隔,确保信息完整。
- 获奖数量统计:优先统计
<li>标签数量(最常见的获奖列表格式),无列表时按换行拆分文本统计,适配不同页面的排版差异。 - 异常兼容:如果页面没有人物信息表格,返回默认的"无数据"和0,避免脚本中断。
内容的提问来源于stack exchange,提问作者George Gkonis
相关产品推荐
相关产品推荐

