如何使用BeautifulSoup提取IMDb剧集页面的演员姓名
解决方案
你之前写的actors = soup.findAll('table',{'cast_list'})仅选中了整个演员表的表格节点,没有深入到存储演员姓名的子标签,所以拿不到目标文本,按以下逻辑修改即可:
- 优先定位到class为
cast_list的演职员表格,推荐使用BeautifulSoup的class_参数匹配class属性,避免和Python内置的class关键字冲突 - 在表格范围内筛选所有指向演员个人页的
<a>标签:这类标签的href属性统一以/name/开头,用这个特征过滤可以避开表格内其他跳转到角色、剧集页的无关链接 - 提取标签文本后调用
.strip()方法,清除文本前后多余的空格、换行符,就能得到干净的演员姓名
可运行参考代码
from bs4 import BeautifulSoup import requests # 请求必须携带合法User-Agent,否则会被IMDb反爬规则拦截,返回403错误页 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36" } resp = requests.get("https://www.imdb.com/title/tt0386676/fullcredits/?ref_=tt_ql_cl", headers=headers) soup = BeautifulSoup(resp.text, "html.parser") actor_names = [] cast_table = soup.find("table", class_="cast_list") # 遍历表格行,跳过第一行表头 for row in cast_table.find_all("tr")[1:]: name_tag = row.find("a", href=lambda attr_val: attr_val and attr_val.startswith("/name/")) if name_tag: actor_names.append(name_tag.text.strip()) # 验证输出,首项即为你示例中提到的Rainn Wilson print(actor_names[:5])
注意:不要直接在表格内全量查找
<a>标签,演员表中混杂了角色页、分集页的跳转链接,不加筛选会提取到大量无关内容,用href特征匹配的准确率更高。
内容的提问来源于stack exchange,提问作者arch200
相关产品推荐
相关产品推荐

