如何用Python从烂番茄电影页面提取编剧的个人页面链接?
如何用Python从烂番茄电影页面提取编剧的个人页面链接?
我来帮你搞定这个问题!之前我也碰到过类似的情况,烂番茄的Movie Info区域元素结构确实有点“撞脸”,很多标签的类名甚至ID都重复,得找精准的定位逻辑才行。
先给你拆解思路,再上实用代码:
第一步:先摸透页面结构
打开浏览器开发者工具(按F12)定位到编剧区块,你会发现烂番茄这类信息一般用**定义列表(<dl>)**组织:<dt>标签放“Screenwriter”这类标签名,<dd>标签里就是对应的编剧链接——这就是我们的突破口!
第二步:用BeautifulSoup精准定位
核心逻辑是:先找到写着“Screenwriter”的<dt>标签,再定位它的下一个兄弟节点<dd>,最后提取里面的<a>链接。
给你写个完整的可运行代码示例:
import requests from bs4 import BeautifulSoup # 目标电影页面URL movie_url = "https://www.rottentomatoes.com/m/dangerous_animals" # 加请求头避免被烂番茄反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } # 获取页面内容 response = requests.get(movie_url, headers=headers) response.raise_for_status() # 请求失败直接抛出异常 soup = BeautifulSoup(response.text, "html.parser") # 定位Screenwriter标签 screenwriter_label = soup.find("dt", string="Screenwriter") if screenwriter_label: # 找到对应的内容容器<dd>,再提取链接 screenwriter_dd = screenwriter_label.find_next_sibling("dd") if screenwriter_dd: screenwriter_a_tag = screenwriter_dd.find("a") if screenwriter_a_tag: # 拼接完整URL(页面里的href是相对路径) full_celebrity_url = f"https://www.rottentomatoes.com{screenwriter_a_tag['href']}" print("提取到的编剧链接:", full_celebrity_url) else: print("未找到编剧信息区块")
万一页面结构变了怎么办?
如果烂番茄哪天改了页面结构,<dt>标签不用了,你可以换个思路:直接定位“Screenwriter”文本节点,再找它的父元素和相邻链接:
# 备选定位方式:先找到文本节点 screenwriter_text = soup.find(string="Screenwriter") if screenwriter_text: # 找到文本的父标签,再找下一个兄弟元素里的链接 link_tag = screenwriter_text.parent.find_next_sibling().find("a") if link_tag: full_url = f"https://www.rottentomatoes.com{link_tag['href']}" print(full_url)
重点提醒
一定要加User-Agent请求头!烂番茄会拦截无标识的爬虫请求,加了之后才能正常获取页面内容。
内容来源于stack exchange
相关产品推荐
相关产品推荐

