You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

基于Python爬取符合条件的IMDB电影预告片的优化方案咨询

Great question! Your current workflow gets the job done, but there are definitely more streamlined ways to pull those trailers without bouncing between IMDb and YouTube. Let me walk you through a few optimized approaches that’ll save you time and reduce friction:

Option 1: Use IMDb/Partner APIs for One-Stop Data

Instead of scraping IMDb first then jumping to YouTube, use an API that lets you fetch both qualifying movie details and their trailer links in a single request. This eliminates the need for cross-platform searching entirely.

  • The official IMDb API (requires a developer account) lets you filter movies by release year, rating, and includes a trailer field that often links directly to the official YouTube trailer or IMDb’s embedded player.
  • If official API access feels like too much hassle, third-party tools like the OMDB API (which pulls data directly from IMDb) work great too. You can filter with parameters for release year and minimum rating, then grab the trailer URL straight from the response.

Here’s a quick code snippet using OMDB (you’ll need a free API key to get started):

import requests

OMDB_API_KEY = "your_api_key_here"
base_url = "http://www.omdbapi.com/"

# Fetch 2024 movies
search_params = {
    "apikey": OMDB_API_KEY,
    "type": "movie",
    "y": "2024",
    "r": "json"
}
search_response = requests.get(base_url, params=search_params)
movies = search_response.json().get("Search", [])

for movie in movies:
    # Get full details to check rating and fetch trailer
    detail_params = {"apikey": OMDB_API_KEY, "i": movie["imdbID"]}
    movie_details = requests.get(base_url, params=detail_params).json()
    
    rating = float(movie_details.get("imdbRating", 0))
    trailer_link = movie_details.get("Trailer")
    
    if rating > 2 and trailer_link:
        # Use yt-dlp to download the trailer (works for YouTube links)
        print(f"Downloading trailer for {movie['Title']}: {trailer_link}")

Option 2: Optimize YouTube Searches with Targeted Queries & APIs

If you still prefer sourcing trailers from YouTube, you can make the search process far more precise (and avoid sifting through random top results) using the YouTube Data API:

  1. When fetching movie info from IMDb, grab the official title and release year to avoid ambiguous results (e.g., "The Flash" could refer to multiple movies).
  2. Use a refined search query like "{movie_title} {release_year} official trailer" and filter results to only show videos from verified channels. This guarantees you get the official trailer immediately.
  3. Pair this with yt-dlp (a maintained fork of youtube-dl) to bulk-download all trailers once you have the video URLs.

Example snippet for targeted YouTube searches:

import requests
import yt_dlp

YOUTUBE_API_KEY = "your_youtube_api_key_here"
youtube_search_url = "https://www.googleapis.com/youtube/v3/search"

def get_official_trailer(movie_title, release_year):
    search_params = {
        "key": YOUTUBE_API_KEY,
        "q": f"{movie_title} {release_year} official trailer",
        "part": "snippet",
        "type": "video",
        "maxResults": 1,
        "videoDefinition": "high"
    }
    response = requests.get(youtube_search_url, params=search_params)
    items = response.json().get("items", [])
    
    if items:
        video_id = items[0]["id"]["videoId"]
        return f"https://www.youtube.com/watch?v={video_id}"
    return None

# Usage with a sample movie from IMDb
movie_title = "Dune: Part Two"
release_year = 2024
trailer_url = get_official_trailer(movie_title, release_year)

if trailer_url:
    # Download the trailer
    ydl_opts = {"outtmpl": f"{movie_title}_trailer.%(ext)s"}
    with yt_dlp.YoutubeDL(ydl_opts) as ydl:
        ydl.download([trailer_url])

If you want to avoid APIs entirely (to skip API keys or rate limits), you can scrape IMDb’s movie pages directly to extract trailer links without ever leaving the platform:

  1. Scrape IMDb’s advanced search page (filtered for 2024 movies with ratings >2) to get your list of target movie URLs.
  2. For each movie page, parse the HTML with BeautifulSoup or Scrapy to find the trailer section (look for elements with data-testid="video-player__play-button" or the "Trailers and Videos" tab).
  3. Extract the underlying YouTube or direct video link, then download with yt-dlp (it handles both IMDb’s internal trailer links and YouTube URLs seamlessly).

Here’s a quick example with BeautifulSoup:

import requests
from bs4 import BeautifulSoup
import yt_dlp

def get_qualifying_movie_links():
    # IMDb advanced search URL for 2024 feature films with rating >2
    search_url = "https://www.imdb.com/search/title/?release_date=2024-01-01,2024-12-31&user_rating=2.0,&title_type=feature"
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36"}
    
    response = requests.get(search_url, headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")
    
    movie_links = []
    for card in soup.find_all("div", class_="ipc-metadata-list-summary-item__tc"):
        link = card.find("a", class_="ipc-title-link-wrapper")["href"]
        movie_links.append(f"https://www.imdb.com{link}")
    
    return movie_links

def extract_trailer_from_imdb_page(movie_url):
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36"}
    response = requests.get(movie_url, headers=headers)
    soup = BeautifulSoup(response.text, "html.parser")
    
    trailer_tag = soup.find("a", {"data-testid": "video-player__play-button"})
    if trailer_tag:
        return f"https://www.imdb.com{trailer_tag['href']}"
    return None

# Run the workflow
movie_links = get_qualifying_movie_links()
for link in movie_links:
    trailer_url = extract_trailer_from_imdb_page(link)
    if trailer_url:
        ydl_opts = {"outtmpl": "%(title)s_trailer.%(ext)s"}
        with yt_dlp.YoutubeDL(ydl_opts) as ydl:
            ydl.download([trailer_url])

Final Recommendation

If you can use an API, go with Option 1—it’s the most stable and least prone to breaking if IMDb changes its page structure. If APIs aren’t an option, Option 3 is better than your original workflow because it keeps everything within IMDb, avoiding the extra YouTube search step.

内容的提问来源于stack exchange,提问作者Barry

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 07:48:32