如何用Python编写函数从archive.is系列短链接提取原始URL?
Hey there! Getting the original URL from archive.is, archive.fo, archive.li, or archive.today links is straightforward with Python. Below are two reliable approaches, with the first being the most stable since it parses the actual archive page content.
Approach 1: Parse the Archive Page with Requests + BeautifulSoup
This method works by fetching the archive page and extracting the original URL directly from the HTML. It's reliable because it targets the visible "source URL" element that archive sites display on every snapshot.
First, install the required packages:
pip install requests beautifulsoup4
Here's the full function with error handling:
import requests from bs4 import BeautifulSoup from urllib.parse import urlparse def get_original_archive_url(archive_url): # Make sure we're dealing with a supported archive domain allowed_domains = {"archive.is", "archive.fo", "archive.li", "archive.today"} parsed_domain = urlparse(archive_url).netloc if parsed_domain not in allowed_domains: raise ValueError("This URL isn't from a supported archive site") try: # Mimic a browser request to avoid being blocked headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } response = requests.get(archive_url, headers=headers, timeout=10) response.raise_for_status() # Catch HTTP errors like 404 or 500 soup = BeautifulSoup(response.text, "html.parser") # First try: Grab the direct source URL link (visible on the archive page) source_link = soup.find("a", class_="source-url") if source_link and source_link.get("href"): return source_link["href"] # Fallback: Check the Open Graph meta tag (some snapshots use this) og_url_meta = soup.find("meta", property="og:url") if og_url_meta and og_url_meta.get("content"): return og_url_meta["content"] # If we can't find the URL anywhere raise ValueError("Couldn't locate the original URL on the archive page") except requests.exceptions.RequestException as e: raise RuntimeError(f"Failed to load the archive page: {str(e)}") from e # Test with your example if __name__ == "__main__": test_link = "http://archive.is/9mIro" try: original_url = get_original_archive_url(test_link) print(f"Original URL: {original_url}") except (ValueError, RuntimeError) as e: print(f"Error: {str(e)}")
How It Works:
- Domain Validation: Checks that the input URL is from one of the supported archive sites to avoid wasted requests.
- Browser Mimicry: Uses a realistic
User-Agentheader to bypass basic anti-scraping measures that archive.is might use. - Primary Extraction: Targets the
<a>tag with thesource-urlclass—this is the link that appears as the original site on the archive page. - Fallback: If the primary method fails, it checks the
og:urlmeta tag, which some archive snapshots use to define the original content URL. - Error Handling: Catches HTTP errors, timeouts, and cases where the original URL can't be found.
Approach 2: Use the Archive Redirect Endpoint (Less Stable)
Some archive.is variants support a redirect endpoint that sends you directly to the original URL. For example, adding /redirect/ before the snapshot hash: http://archive.is/redirect/9mIro. However, this endpoint isn't consistently supported across all archive domains, so it's less reliable than parsing the page.
Here's a quick example of this method:
import requests def get_original_via_redirect(archive_url): redirect_url = archive_url.replace("archive.is/", "archive.is/redirect/") try: # Use HEAD request to avoid downloading the entire page response = requests.head(redirect_url, allow_redirects=True, timeout=10) return response.url except requests.exceptions.RequestException as e: raise RuntimeError(f"Redirect failed: {str(e)}") from e
Note: This method can break if archive sites change their redirect structure, so stick with Approach 1 for long-term reliability.
内容的提问来源于stack exchange,提问作者Joshua Meyers

