如何用Pandas与BeautifulSoup抓取演员姓名及详情页生日信息
解决方案:提取演员姓名列表并获取生日信息
1. 将演员姓名存入列表
在你已有的代码基础上,只需初始化一个空列表,每次提取到演员姓名时将其追加到列表中即可:
- 初始化空列表:
actor_names = [] - 提取到姓名后,执行
actor_names.append(actor_name)完成存储
2. 从演员详情页获取生日信息
要获取生日,需先从原页面提取演员的维基百科详情页链接,再访问该链接并解析页面中的生日字段(维基百科通常用 bday 类标记生日)。
完整可运行代码
import requests from bs4 import BeautifulSoup import pandas as pd # 目标页面URL target_url = "https://en.wikipedia.org/w/index.php?title=Chernobyl_(miniseries)&direction=next&oldid=899952822" # 访问切尔诺贝利迷你剧页面,提取前三位演员信息 response = requests.get(target_url) soup = BeautifulSoup(response.text, "html.parser") # 定位演员列表所在的表格 cast_heading = soup.find("span", id="Cast") cast_table = cast_heading.find_next("table", class_="wikitable") # 获取前三位演员的行(跳过表头行) actor_rows = cast_table.find_all("tr")[1:4] # 初始化存储容器 actor_names = [] actor_birthday_data = [] # 遍历每位演员,获取姓名和生日 for row in actor_rows: # 提取演员姓名和详情页链接 actor_link = row.find("td").find("a") actor_name = actor_link.text.strip() actor_names.append(actor_name) # 拼接演员详情页完整URL actor_page_url = "https://en.wikipedia.org" + actor_link["href"] # 访问演员详情页并解析生日 actor_response = requests.get(actor_page_url) actor_soup = BeautifulSoup(actor_response.text, "html.parser") birthday = actor_soup.find("span", class_="bday").text.strip() if actor_soup.find("span", class_="bday") else "未找到" # 存储姓名和生日的对应关系 actor_birthday_data.append({"姓名": actor_name, "生日": birthday}) # 用Pandas整理数据(可选) actor_df = pd.DataFrame(actor_birthday_data) print("演员姓名列表:", actor_names) print("\n演员生日信息:") print(actor_df)
关键说明
- 姓名存储:最终
actor_names列表会包含['Jared Harris', 'Stellan Skarsgård', 'Paul Ritter'],可直接用于后续逻辑。 - 生日提取:通过定位
bday类的元素获取标准格式的生日,同时处理了字段不存在的异常场景。 - Pandas整合:将数据转换为DataFrame后,支持导出为CSV、Excel等格式,便于后续分析。
内容的提问来源于stack exchange,提问作者maja kralj
相关产品推荐
相关产品推荐

