You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取进阶:如何跟踪链接获取IMSDb恐怖电影剧本

如何跟踪IMSDb的链接获取恐怖电影剧本

完全可以用requests.get()来跟踪链接,每一次页面跳转本质都是发起新的HTTP请求。你已经拿到了第一阶段的电影链接,接下来只需要对每个电影链接发起请求,找到剧本入口,再对剧本页发起请求提取内容即可。

代码扩展步骤

1. 补全电影链接的完整URL

你现在拿到的lista里的链接是相对路径(比如/Movie Scripts/Alien Script.html),需要拼接成完整URL才能用requests.get()请求:

base_url = "https://imsdb.com"
full_movie_links = [base_url + link for link in lista]

2. 遍历电影链接,获取剧本入口

对每个完整的电影页面URL发起请求,用BeautifulSoup找到剧本入口的链接。通常剧本入口是带有"Read XXXX Script"字样的a标签:

# 加headers模拟浏览器请求,避免被反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

script_links = []
for movie_url in full_movie_links:
    # 请求电影页面
    movie_response = requests.get(movie_url, headers=headers)
    movie_soup = BeautifulSoup(movie_response.text, 'lxml')
    
    # 定位剧本入口链接
    script_a_tag = movie_soup.find('a', string=lambda text: text and "Read" in text and "Script" in text)
    if script_a_tag:
        script_href = script_a_tag.get('href')
        full_script_url = base_url + script_href
        script_links.append(full_script_url)
    else:
        print(f"未找到{movie_url}的剧本入口")

3. 请求剧本页面,提取剧本内容

拿到剧本页的完整URL后,再次发起请求,IMSDb的剧本通常放在<pre>标签里,直接提取即可:

scripts = []
for script_url in script_links:
    script_response = requests.get(script_url, headers=headers)
    script_soup = BeautifulSoup(script_response.text, 'lxml')
    
    # 提取剧本内容
    script_content = script_soup.find('pre').text
    scripts.append(script_content)
    
    # 保存为本地文件(Colab左侧文件面板可查看下载)
    movie_name = script_url.split('/')[-1].replace(' Script.html', '')
    with open(f"{movie_name}_script.txt", "w", encoding="utf-8") as f:
        f.write(script_content)

完整整合代码

把上面的步骤和你的初始代码整合,同时加上异常处理避免报错中断:

import requests
from bs4 import BeautifulSoup
import lxml
import time

# 初始获取电影列表部分
website = 'https://imsdb.com/genre/Horror'
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}
resultado = requests.get(website, headers=headers)
contenido = resultado.text

soup = BeautifulSoup(contenido, 'lxml')
info = soup.find_all('td', {'valign':'top'})
info2 = info[-1]

lista = []
for link in info2.find_all('a'):
    aux = link.get('href')
    lista.append(aux)

# 补全电影完整链接
base_url = "https://imsdb.com"
full_movie_links = [base_url + link for link in lista]

# 获取剧本入口链接
script_links = []
for movie_url in full_movie_links:
    try:
        movie_response = requests.get(movie_url, headers=headers)
        movie_response.raise_for_status()  # 检查请求是否成功
        movie_soup = BeautifulSoup(movie_response.text, 'lxml')
        
        script_a_tag = movie_soup.find('a', string=lambda t: t and "Read" in t and "Script" in t)
        if script_a_tag:
            script_href = script_a_tag.get('href')
            full_script_url = base_url + script_href
            script_links.append(full_script_url)
        else:
            print(f"{movie_url} 未找到剧本入口")
        time.sleep(1)  # 加延迟避免请求太频繁
    except Exception as e:
        print(f"请求{movie_url}失败: {str(e)}")

# 提取并保存剧本
for idx, script_url in enumerate(script_links):
    try:
        script_response = requests.get(script_url, headers=headers)
        script_response.raise_for_status()
        script_soup = BeautifulSoup(script_response.text, 'lxml')
        
        script_pre = script_soup.find('pre')
        if script_pre:
            script_content = script_pre.text
            movie_name = script_url.split('/')[-1].replace('.html', '').replace(' Script', '')
            with open(f"{idx+1}_{movie_name}_script.txt", "w", encoding="utf-8") as f:
                f.write(script_content)
            print(f"已保存 {movie_name} 的剧本")
        else:
            print(f"{script_url} 未找到剧本内容")
        time.sleep(1)
    except Exception as e:
        print(f"请求{script_url}失败: {str(e)}")

注意事项

  • 页面结构可能变化:如果后续找不到元素,打开对应页面用浏览器开发者工具查看元素的class或标签,调整BeautifulSoup的选择器。
  • 反爬限制:保持time.sleep()的延迟,避免短时间内大量请求被封禁IP。
  • Colab文件管理:保存的txt文件可在Colab左侧"文件"面板找到,右键即可下载。

内容的提问来源于stack exchange,提问作者Zarcklet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 07:00:57