基于Python爬取符合条件的IMDB电影预告片的优化方案咨询
Great question! Your current workflow gets the job done, but there are definitely more streamlined ways to pull those trailers without bouncing between IMDb and YouTube. Let me walk you through a few optimized approaches that’ll save you time and reduce friction:
Option 1: Use IMDb/Partner APIs for One-Stop Data
Instead of scraping IMDb first then jumping to YouTube, use an API that lets you fetch both qualifying movie details and their trailer links in a single request. This eliminates the need for cross-platform searching entirely.
- The official IMDb API (requires a developer account) lets you filter movies by release year, rating, and includes a
trailerfield that often links directly to the official YouTube trailer or IMDb’s embedded player. - If official API access feels like too much hassle, third-party tools like the OMDB API (which pulls data directly from IMDb) work great too. You can filter with parameters for release year and minimum rating, then grab the trailer URL straight from the response.
Here’s a quick code snippet using OMDB (you’ll need a free API key to get started):
import requests OMDB_API_KEY = "your_api_key_here" base_url = "http://www.omdbapi.com/" # Fetch 2024 movies search_params = { "apikey": OMDB_API_KEY, "type": "movie", "y": "2024", "r": "json" } search_response = requests.get(base_url, params=search_params) movies = search_response.json().get("Search", []) for movie in movies: # Get full details to check rating and fetch trailer detail_params = {"apikey": OMDB_API_KEY, "i": movie["imdbID"]} movie_details = requests.get(base_url, params=detail_params).json() rating = float(movie_details.get("imdbRating", 0)) trailer_link = movie_details.get("Trailer") if rating > 2 and trailer_link: # Use yt-dlp to download the trailer (works for YouTube links) print(f"Downloading trailer for {movie['Title']}: {trailer_link}")
Option 2: Optimize YouTube Searches with Targeted Queries & APIs
If you still prefer sourcing trailers from YouTube, you can make the search process far more precise (and avoid sifting through random top results) using the YouTube Data API:
- When fetching movie info from IMDb, grab the official title and release year to avoid ambiguous results (e.g., "The Flash" could refer to multiple movies).
- Use a refined search query like
"{movie_title} {release_year} official trailer"and filter results to only show videos from verified channels. This guarantees you get the official trailer immediately. - Pair this with
yt-dlp(a maintained fork of youtube-dl) to bulk-download all trailers once you have the video URLs.
Example snippet for targeted YouTube searches:
import requests import yt_dlp YOUTUBE_API_KEY = "your_youtube_api_key_here" youtube_search_url = "https://www.googleapis.com/youtube/v3/search" def get_official_trailer(movie_title, release_year): search_params = { "key": YOUTUBE_API_KEY, "q": f"{movie_title} {release_year} official trailer", "part": "snippet", "type": "video", "maxResults": 1, "videoDefinition": "high" } response = requests.get(youtube_search_url, params=search_params) items = response.json().get("items", []) if items: video_id = items[0]["id"]["videoId"] return f"https://www.youtube.com/watch?v={video_id}" return None # Usage with a sample movie from IMDb movie_title = "Dune: Part Two" release_year = 2024 trailer_url = get_official_trailer(movie_title, release_year) if trailer_url: # Download the trailer ydl_opts = {"outtmpl": f"{movie_title}_trailer.%(ext)s"} with yt_dlp.YoutubeDL(ydl_opts) as ydl: ydl.download([trailer_url])
Option 3: Scrape IMDb Directly for Trailer Links
If you want to avoid APIs entirely (to skip API keys or rate limits), you can scrape IMDb’s movie pages directly to extract trailer links without ever leaving the platform:
- Scrape IMDb’s advanced search page (filtered for 2024 movies with ratings >2) to get your list of target movie URLs.
- For each movie page, parse the HTML with BeautifulSoup or Scrapy to find the trailer section (look for elements with
data-testid="video-player__play-button"or the "Trailers and Videos" tab). - Extract the underlying YouTube or direct video link, then download with
yt-dlp(it handles both IMDb’s internal trailer links and YouTube URLs seamlessly).
Here’s a quick example with BeautifulSoup:
import requests from bs4 import BeautifulSoup import yt_dlp def get_qualifying_movie_links(): # IMDb advanced search URL for 2024 feature films with rating >2 search_url = "https://www.imdb.com/search/title/?release_date=2024-01-01,2024-12-31&user_rating=2.0,&title_type=feature" headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36"} response = requests.get(search_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") movie_links = [] for card in soup.find_all("div", class_="ipc-metadata-list-summary-item__tc"): link = card.find("a", class_="ipc-title-link-wrapper")["href"] movie_links.append(f"https://www.imdb.com{link}") return movie_links def extract_trailer_from_imdb_page(movie_url): headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/123.0.0.0 Safari/537.36"} response = requests.get(movie_url, headers=headers) soup = BeautifulSoup(response.text, "html.parser") trailer_tag = soup.find("a", {"data-testid": "video-player__play-button"}) if trailer_tag: return f"https://www.imdb.com{trailer_tag['href']}" return None # Run the workflow movie_links = get_qualifying_movie_links() for link in movie_links: trailer_url = extract_trailer_from_imdb_page(link) if trailer_url: ydl_opts = {"outtmpl": "%(title)s_trailer.%(ext)s"} with yt_dlp.YoutubeDL(ydl_opts) as ydl: ydl.download([trailer_url])
Final Recommendation
If you can use an API, go with Option 1—it’s the most stable and least prone to breaking if IMDb changes its page structure. If APIs aren’t an option, Option 3 is better than your original workflow because it keeps everything within IMDb, avoiding the extra YouTube search step.
内容的提问来源于stack exchange,提问作者Barry

