You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何批量提取imsdb.com所有电影脚本并导出为Word文件

IMSDB电影脚本批量提取落地方案

前置依赖安装

先安装需要的第三方库:

pip install beautifulsoup4 requests python-docx

实现步骤

步骤1:提取全量脚本URL

从索引页抓取所有符合/scripts/*.html格式的有效链接,自动去重、过滤无效地址:

import requests
from bs4 import BeautifulSoup
import time
from docx import Document
import os

INDEX_URL = "https://imsdb.com/all-scripts.html"
# 伪装浏览器请求头,避免被反爬机制拦截
HEADERS = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

def get_all_script_urls():
    res = requests.get(INDEX_URL, headers=HEADERS)
    res.encoding = "utf-8"
    soup = BeautifulSoup(res.text, "html.parser")
    script_urls = set() # 用set结构自动去重
    for a_tag in soup.find_all("a", href=True):
        href = a_tag["href"]
        if href.startswith("/scripts/") and href.endswith(".html"):
            full_url = f"https://imsdb.com{href}"
            script_urls.add(full_url)
    return list(script_urls)

步骤2:优化单脚本内容提取逻辑

原全页提取逻辑会拿到导航、广告、底部信息等冗余内容,可直接定位脚本正文节点提升提取精度:

def get_single_script_content(script_url):
    try:
        res = requests.get(script_url, headers=HEADERS, timeout=10)
        res.encoding = "utf-8"
        soup = BeautifulSoup(res.text, "html.parser")
        # 站点正文统一存在class为scrtext的td节点或者pre节点中
        content_node = soup.find("td", class_="scrtext")
        if not content_node:
            content_node = soup.find("pre")
        if content_node:
            return content_node.get_text().strip()
        else:
            # 特殊页面降级到全页提取
            return soup.get_text().strip()
    except Exception as e:
        print(f"抓取{script_url}失败:{str(e)}")
        return ""

步骤3:批量导出Word文件

两种导出模式按需选择:

模式1:每个脚本单独生成一个Word文件

def export_to_separate_docs(script_urls, save_dir="imsdb_scripts"):
    os.makedirs(save_dir, exist_ok=True)
    for idx, url in enumerate(script_urls):
        # 从URL中提取脚本名称作为文件名
        script_name = url.split("/")[-1].replace(".html", "").replace("-", " ")
        print(f"正在处理第{idx+1}/{len(script_urls)}个脚本:{script_name}")
        content = get_single_script_content(url)
        if not content:
            continue
        doc = Document()
        doc.add_heading(script_name, level=1)
        doc.add_paragraph(content)
        doc.save(os.path.join(save_dir, f"{script_name}.docx"))
        # 增加请求间隔,避免请求过快被封IP,建议2-5秒
        time.sleep(2)

模式2:所有脚本合并到同一个Word文件

def export_to_single_doc(script_urls, save_path="imsdb_all_scripts.docx"):
    doc = Document()
    doc.add_heading("IMSDB全量电影脚本", level=0)
    for idx, url in enumerate(script_urls):
        script_name = url.split("/")[-1].replace(".html", "").replace("-", " ")
        print(f"正在处理第{idx+1}/{len(script_urls)}个脚本:{script_name}")
        content = get_single_script_content(url)
        if not content:
            continue
        doc.add_heading(script_name, level=1)
        doc.add_paragraph(content)
        # 加分页符分隔不同脚本
        doc.add_page_break()
        time.sleep(2)
    doc.save(save_path)

调用示例

if __name__ == "__main__":
    all_urls = get_all_script_urls()
    print(f"共获取到{len(all_urls)}个脚本地址")
    # 按需选择对应导出模式执行即可
    # export_to_separate_docs(all_urls)
    # export_to_single_doc(all_urls)

注意事项

  • 可根据自身网络情况调整time.sleep的延时时长,站点有反爬机制,请求过快会触发临时IP封禁
  • 如果抓取中途中断,可以把已经处理完成的脚本名称存储到本地列表,下次启动时跳过已处理内容,无需从头开始抓取
  • 少量老脚本页面结构特殊,如果提取到空内容,可以单独针对这类地址调整DOM定位逻辑

内容的提问来源于stack exchange,提问作者Zhomart

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.10.04 02:12:02