如何用Python抓取Mega.nz存档页面中的文本内容?
Ah, I’ve dealt with this exact frustration with Mega.nz—their heavy use of client-side JavaScript means all the folder/file names you see are rendered after the initial HTML loads, so static scraping tools that just fetch raw HTML will come up empty. Here are two solid, practical approaches to solve this efficiently:
1. Use a Headless Browser to Simulate Real User Navigation
This is the most straightforward "set it and forget it" method if you don’t want to dive into API details. Tools like Playwright or Selenium spin up an invisible browser, let the page fully render all JS content, then let you extract the text you need.
Example with Playwright (Lightweight & Reliable)
First, install the package and browser binary:
pip install playwright playwright install chromium
Then write a script to handle your links:
from playwright.sync_api import sync_playwright def scrape_mega_folder_name(url): with sync_playwright() as p: # Launch a headless browser (add `headless=False` to see the browser window) browser = p.chromium.launch(headless=True, args=["--disable-blink-features=AutomationControlled"]) page = browser.new_page(user_agent="Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36") try: page.goto(url, wait_until="networkidle") # Wait for the folder name element to load (adjust the selector if needed) folder_name_selector = ".header-title h1" page.wait_for_selector(folder_name_selector, timeout=10000) folder_name = page.locator(folder_name_selector).text_content() return folder_name.strip() finally: browser.close() # Test it with your link mega_url = "https://mega.nz/folder/your-folder-id#your-key" print(scrape_mega_folder_name(mega_url))
Pro Tip: If you hit anti-scraping detection, tweak the args and user_agent values to make the browser look more like a real human user.
2. Use Mega’s API (Faster for Large Volumes)
For handling hundreds or thousands of links, simulating a browser for each one is slow and resource-heavy. Instead, use Mega’s public API directly—this lets you fetch folder metadata without rendering any HTML at all.
We’ll use the mega.py third-party library (built specifically for interacting with Mega’s API):
pip install mega.py
Then script to pull folder info:
from mega import Mega def get_mega_folder_details(url): mega = Mega() # Fetch the public folder object directly from the URL folder = mega.get_public_folder(url) # Extract folder name and optional content list folder_name = folder['name'] folder_contents = [node['name'] for node in folder['nodes'] if node['type'] != 1] # Filter out subfolders if needed return { "folder_name": folder_name, "contents": folder_contents } # Test with your link mega_url = "https://mega.nz/folder/your-folder-id#your-key" details = get_mega_folder_details(mega_url) print(f"Folder Name: {details['folder_name']}") print("Contents:") for item in details['contents']: print(f"- {item}")
Note: This method works for public links only (no login required). If your links need authentication, add a login step with mega.login(email, password).
Which One to Choose?
- Use the headless browser method if you’re new to scraping or need to extract other dynamic elements (like descriptions) alongside names.
- Use the API method if you’re dealing with a large number of links—it’s faster, uses less CPU/memory, and avoids browser detection headaches.
内容的提问来源于stack exchange,提问作者EllipticalInitial

