网页爬取进阶:如何跟踪链接获取IMSDb恐怖电影剧本
如何跟踪IMSDb的链接获取恐怖电影剧本
完全可以用requests.get()来跟踪链接,每一次页面跳转本质都是发起新的HTTP请求。你已经拿到了第一阶段的电影链接,接下来只需要对每个电影链接发起请求,找到剧本入口,再对剧本页发起请求提取内容即可。
代码扩展步骤
1. 补全电影链接的完整URL
你现在拿到的lista里的链接是相对路径(比如/Movie Scripts/Alien Script.html),需要拼接成完整URL才能用requests.get()请求:
base_url = "https://imsdb.com" full_movie_links = [base_url + link for link in lista]
2. 遍历电影链接,获取剧本入口
对每个完整的电影页面URL发起请求,用BeautifulSoup找到剧本入口的链接。通常剧本入口是带有"Read XXXX Script"字样的a标签:
# 加headers模拟浏览器请求,避免被反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } script_links = [] for movie_url in full_movie_links: # 请求电影页面 movie_response = requests.get(movie_url, headers=headers) movie_soup = BeautifulSoup(movie_response.text, 'lxml') # 定位剧本入口链接 script_a_tag = movie_soup.find('a', string=lambda text: text and "Read" in text and "Script" in text) if script_a_tag: script_href = script_a_tag.get('href') full_script_url = base_url + script_href script_links.append(full_script_url) else: print(f"未找到{movie_url}的剧本入口")
3. 请求剧本页面,提取剧本内容
拿到剧本页的完整URL后,再次发起请求,IMSDb的剧本通常放在<pre>标签里,直接提取即可:
scripts = [] for script_url in script_links: script_response = requests.get(script_url, headers=headers) script_soup = BeautifulSoup(script_response.text, 'lxml') # 提取剧本内容 script_content = script_soup.find('pre').text scripts.append(script_content) # 保存为本地文件(Colab左侧文件面板可查看下载) movie_name = script_url.split('/')[-1].replace(' Script.html', '') with open(f"{movie_name}_script.txt", "w", encoding="utf-8") as f: f.write(script_content)
完整整合代码
把上面的步骤和你的初始代码整合,同时加上异常处理避免报错中断:
import requests from bs4 import BeautifulSoup import lxml import time # 初始获取电影列表部分 website = 'https://imsdb.com/genre/Horror' headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } resultado = requests.get(website, headers=headers) contenido = resultado.text soup = BeautifulSoup(contenido, 'lxml') info = soup.find_all('td', {'valign':'top'}) info2 = info[-1] lista = [] for link in info2.find_all('a'): aux = link.get('href') lista.append(aux) # 补全电影完整链接 base_url = "https://imsdb.com" full_movie_links = [base_url + link for link in lista] # 获取剧本入口链接 script_links = [] for movie_url in full_movie_links: try: movie_response = requests.get(movie_url, headers=headers) movie_response.raise_for_status() # 检查请求是否成功 movie_soup = BeautifulSoup(movie_response.text, 'lxml') script_a_tag = movie_soup.find('a', string=lambda t: t and "Read" in t and "Script" in t) if script_a_tag: script_href = script_a_tag.get('href') full_script_url = base_url + script_href script_links.append(full_script_url) else: print(f"{movie_url} 未找到剧本入口") time.sleep(1) # 加延迟避免请求太频繁 except Exception as e: print(f"请求{movie_url}失败: {str(e)}") # 提取并保存剧本 for idx, script_url in enumerate(script_links): try: script_response = requests.get(script_url, headers=headers) script_response.raise_for_status() script_soup = BeautifulSoup(script_response.text, 'lxml') script_pre = script_soup.find('pre') if script_pre: script_content = script_pre.text movie_name = script_url.split('/')[-1].replace('.html', '').replace(' Script', '') with open(f"{idx+1}_{movie_name}_script.txt", "w", encoding="utf-8") as f: f.write(script_content) print(f"已保存 {movie_name} 的剧本") else: print(f"{script_url} 未找到剧本内容") time.sleep(1) except Exception as e: print(f"请求{script_url}失败: {str(e)}")
注意事项
- 页面结构可能变化:如果后续找不到元素,打开对应页面用浏览器开发者工具查看元素的class或标签,调整BeautifulSoup的选择器。
- 反爬限制:保持
time.sleep()的延迟,避免短时间内大量请求被封禁IP。 - Colab文件管理:保存的txt文件可在Colab左侧"文件"面板找到,右键即可下载。
内容的提问来源于stack exchange,提问作者Zarcklet
相关产品推荐
相关产品推荐

