You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

问询:基于Selenium的社交媒体图片视频下载Python脚本及实现路线图

Roadmap to Build a Python Social Media Media Downloader

Got it, building a Python script to download images and videos from social media URLs is totally feasible. Here’s a step-by-step roadmap to help you implement it, complete with key code snippets and important considerations:

1. Set Up Command Line Argument Parsing

First, you need to accept the target URL and save path as command-line inputs. The argparse module is perfect for this—it handles validation and help text out of the box.

import argparse
import os

def parse_args():
    parser = argparse.ArgumentParser(description='Download images and videos from a social media URL.')
    parser.add_argument('url', help='Target social media URL (e.g., Instagram post, TikTok video)')
    parser.add_argument('path', help='Directory to save downloaded media files')
    args = parser.parse_args()
    
    # Create the save directory if it doesn't exist
    if not os.path.isdir(args.path):
        os.makedirs(args.path)
    
    return args.url, args.path

2. Choose a Content Extraction Tool

Social media sites vary in how they serve content:

  • Static sites: Content loads in the initial HTML (rare for modern social media). Use requests + BeautifulSoup for this.
  • Dynamic sites: Content loads via JavaScript (most common). Use selenium or playwright to render the full page before scraping.

⚠️ Important: Always check the site’s Terms of Service before scraping—many platforms prohibit automated downloading of their content.

3. Extract Media URLs from the Page

Static Content Example (requests + BeautifulSoup)

For sites where media URLs are in the initial HTML:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin

def extract_media_static(base_url):
    # Spoof a browser user-agent to avoid being blocked
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}
    response = requests.get(base_url, headers=headers)
    response.raise_for_status()  # Raise error if request fails
    
    soup = BeautifulSoup(response.text, 'html.parser')
    media_urls = []
    
    # Extract image URLs (check common attributes like src and data-src)
    for img in soup.find_all('img'):
        img_url = img.get('src') or img.get('data-src')
        if img_url:
            full_url = urljoin(base_url, img_url)
            media_urls.append(full_url)
    
    # Extract video URLs
    for video in soup.find_all('video'):
        video_url = video.get('src')
        if video_url:
            full_url = urljoin(base_url, video_url)
            media_urls.append(full_url)
    
    return list(set(media_urls))  # Remove duplicates

Dynamic Content Example (selenium)

For sites that load media via JavaScript:

from selenium import webdriver
from selenium.webdriver.common.by import By
import time

def extract_media_dynamic(base_url):
    # Initialize Chrome driver (ensure ChromeDriver is installed and in your PATH)
    options = webdriver.ChromeOptions()
    options.add_argument('--headless=new')  # Run in background without a window
    driver = webdriver.Chrome(options=options)
    
    driver.get(base_url)
    time.sleep(3)  # Wait for JS content to load (adjust based on site speed)
    
    media_urls = []
    
    # Extract images
    imgs = driver.find_elements(By.TAG_NAME, 'img')
    for img in imgs:
        img_url = img.get_attribute('src') or img.get_attribute('data-src')
        if img_url:
            media_urls.append(img_url)
    
    # Extract videos
    videos = driver.find_elements(By.TAG_NAME, 'video')
    for video in videos:
        video_url = video.get_attribute('src')
        if video_url:
            media_urls.append(video_url)
    
    driver.quit()
    return list(set(media_urls))

4. Download and Save Media Files

Write a function to download each media URL while preserving the original file format:

import requests

def download_media(media_url, save_path):
    try:
        headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'}
        # Stream content to handle large files without using too much memory
        response = requests.get(media_url, headers=headers, stream=True)
        response.raise_for_status()
        
        # Extract filename from URL (clean up query parameters if present)
        filename = media_url.split('/')[-1].split('?')[0]
        full_save_path = os.path.join(save_path, filename)
        
        # Write content to file
        with open(full_save_path, 'wb') as f:
            for chunk in response.iter_content(chunk_size=8192):
                f.write(chunk)
        
        print(f"Successfully saved: {full_save_path}")
    except Exception as e:
        print(f"Failed to download {media_url}: {str(e)}")

5. Combine All Components into a Main Function

Tie everything together into a runnable script:

def main():
    url, save_path = parse_args()
    
    # Try static extraction first, fall back to dynamic if it fails
    try:
        media_urls = extract_media_static(url)
    except Exception as static_error:
        print(f"Static extraction failed: {static_error}. Trying dynamic extraction...")
        media_urls = extract_media_dynamic(url)
    
    # Filter out non-image/video URLs (optional, based on file extensions)
    allowed_extensions = ('.jpg', '.jpeg', '.png', '.gif', '.mp4', '.mov', '.webm')
    filtered_media = [url for url in media_urls if url.lower().endswith(allowed_extensions)]
    
    if not filtered_media:
        print("No images or videos found at the provided URL.")
        return
    
    # Download all valid media files
    for media_url in filtered_media:
        download_media(media_url, save_path)

if __name__ == "__main__":
    main()

6. Enhancements & Best Practices

  • Rate Limiting: Add time.sleep(1) between downloads to avoid overwhelming the server.
  • Media Type Validation: Check the Content-Type header of the media URL to confirm it’s an image or video, instead of relying solely on extensions.
  • Handle Embedded Media: For links to external platforms (e.g., YouTube), use specialized libraries like pytube to extract the video URL.
  • Error Handling: Add more specific exception handling for network issues, permission errors, and invalid URLs.

内容的提问来源于stack exchange,提问作者stack

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.06 19:37:54