问询:基于Selenium的社交媒体图片视频下载Python脚本及实现路线图
Got it, building a Python script to download images and videos from social media URLs is totally feasible. Here’s a step-by-step roadmap to help you implement it, complete with key code snippets and important considerations:
1. Set Up Command Line Argument Parsing
First, you need to accept the target URL and save path as command-line inputs. The argparse module is perfect for this—it handles validation and help text out of the box.
import argparse import os def parse_args(): parser = argparse.ArgumentParser(description='Download images and videos from a social media URL.') parser.add_argument('url', help='Target social media URL (e.g., Instagram post, TikTok video)') parser.add_argument('path', help='Directory to save downloaded media files') args = parser.parse_args() # Create the save directory if it doesn't exist if not os.path.isdir(args.path): os.makedirs(args.path) return args.url, args.path
2. Choose a Content Extraction Tool
Social media sites vary in how they serve content:
- Static sites: Content loads in the initial HTML (rare for modern social media). Use
requests+BeautifulSoupfor this. - Dynamic sites: Content loads via JavaScript (most common). Use
seleniumorplaywrightto render the full page before scraping.
⚠️ Important: Always check the site’s Terms of Service before scraping—many platforms prohibit automated downloading of their content.
3. Extract Media URLs from the Page
Static Content Example (requests + BeautifulSoup)
For sites where media URLs are in the initial HTML:
import requests from bs4 import BeautifulSoup from urllib.parse import urljoin def extract_media_static(base_url): # Spoof a browser user-agent to avoid being blocked headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'} response = requests.get(base_url, headers=headers) response.raise_for_status() # Raise error if request fails soup = BeautifulSoup(response.text, 'html.parser') media_urls = [] # Extract image URLs (check common attributes like src and data-src) for img in soup.find_all('img'): img_url = img.get('src') or img.get('data-src') if img_url: full_url = urljoin(base_url, img_url) media_urls.append(full_url) # Extract video URLs for video in soup.find_all('video'): video_url = video.get('src') if video_url: full_url = urljoin(base_url, video_url) media_urls.append(full_url) return list(set(media_urls)) # Remove duplicates
Dynamic Content Example (selenium)
For sites that load media via JavaScript:
from selenium import webdriver from selenium.webdriver.common.by import By import time def extract_media_dynamic(base_url): # Initialize Chrome driver (ensure ChromeDriver is installed and in your PATH) options = webdriver.ChromeOptions() options.add_argument('--headless=new') # Run in background without a window driver = webdriver.Chrome(options=options) driver.get(base_url) time.sleep(3) # Wait for JS content to load (adjust based on site speed) media_urls = [] # Extract images imgs = driver.find_elements(By.TAG_NAME, 'img') for img in imgs: img_url = img.get_attribute('src') or img.get_attribute('data-src') if img_url: media_urls.append(img_url) # Extract videos videos = driver.find_elements(By.TAG_NAME, 'video') for video in videos: video_url = video.get_attribute('src') if video_url: media_urls.append(video_url) driver.quit() return list(set(media_urls))
4. Download and Save Media Files
Write a function to download each media URL while preserving the original file format:
import requests def download_media(media_url, save_path): try: headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'} # Stream content to handle large files without using too much memory response = requests.get(media_url, headers=headers, stream=True) response.raise_for_status() # Extract filename from URL (clean up query parameters if present) filename = media_url.split('/')[-1].split('?')[0] full_save_path = os.path.join(save_path, filename) # Write content to file with open(full_save_path, 'wb') as f: for chunk in response.iter_content(chunk_size=8192): f.write(chunk) print(f"Successfully saved: {full_save_path}") except Exception as e: print(f"Failed to download {media_url}: {str(e)}")
5. Combine All Components into a Main Function
Tie everything together into a runnable script:
def main(): url, save_path = parse_args() # Try static extraction first, fall back to dynamic if it fails try: media_urls = extract_media_static(url) except Exception as static_error: print(f"Static extraction failed: {static_error}. Trying dynamic extraction...") media_urls = extract_media_dynamic(url) # Filter out non-image/video URLs (optional, based on file extensions) allowed_extensions = ('.jpg', '.jpeg', '.png', '.gif', '.mp4', '.mov', '.webm') filtered_media = [url for url in media_urls if url.lower().endswith(allowed_extensions)] if not filtered_media: print("No images or videos found at the provided URL.") return # Download all valid media files for media_url in filtered_media: download_media(media_url, save_path) if __name__ == "__main__": main()
6. Enhancements & Best Practices
- Rate Limiting: Add
time.sleep(1)between downloads to avoid overwhelming the server. - Media Type Validation: Check the
Content-Typeheader of the media URL to confirm it’s an image or video, instead of relying solely on extensions. - Handle Embedded Media: For links to external platforms (e.g., YouTube), use specialized libraries like
pytubeto extract the video URL. - Error Handling: Add more specific exception handling for network issues, permission errors, and invalid URLs.
内容的提问来源于stack exchange,提问作者stack

