You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python编写函数从archive.is系列短链接提取原始URL?

Hey there! Getting the original URL from archive.is, archive.fo, archive.li, or archive.today links is straightforward with Python. Below are two reliable approaches, with the first being the most stable since it parses the actual archive page content.

Approach 1: Parse the Archive Page with Requests + BeautifulSoup

This method works by fetching the archive page and extracting the original URL directly from the HTML. It's reliable because it targets the visible "source URL" element that archive sites display on every snapshot.

First, install the required packages:

pip install requests beautifulsoup4

Here's the full function with error handling:

import requests
from bs4 import BeautifulSoup
from urllib.parse import urlparse

def get_original_archive_url(archive_url):
    # Make sure we're dealing with a supported archive domain
    allowed_domains = {"archive.is", "archive.fo", "archive.li", "archive.today"}
    parsed_domain = urlparse(archive_url).netloc
    
    if parsed_domain not in allowed_domains:
        raise ValueError("This URL isn't from a supported archive site")
    
    try:
        # Mimic a browser request to avoid being blocked
        headers = {
            "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
        }
        response = requests.get(archive_url, headers=headers, timeout=10)
        response.raise_for_status()  # Catch HTTP errors like 404 or 500
        
        soup = BeautifulSoup(response.text, "html.parser")
        
        # First try: Grab the direct source URL link (visible on the archive page)
        source_link = soup.find("a", class_="source-url")
        if source_link and source_link.get("href"):
            return source_link["href"]
        
        # Fallback: Check the Open Graph meta tag (some snapshots use this)
        og_url_meta = soup.find("meta", property="og:url")
        if og_url_meta and og_url_meta.get("content"):
            return og_url_meta["content"]
        
        # If we can't find the URL anywhere
        raise ValueError("Couldn't locate the original URL on the archive page")
    
    except requests.exceptions.RequestException as e:
        raise RuntimeError(f"Failed to load the archive page: {str(e)}") from e

# Test with your example
if __name__ == "__main__":
    test_link = "http://archive.is/9mIro"
    try:
        original_url = get_original_archive_url(test_link)
        print(f"Original URL: {original_url}")
    except (ValueError, RuntimeError) as e:
        print(f"Error: {str(e)}")

How It Works:

  • Domain Validation: Checks that the input URL is from one of the supported archive sites to avoid wasted requests.
  • Browser Mimicry: Uses a realistic User-Agent header to bypass basic anti-scraping measures that archive.is might use.
  • Primary Extraction: Targets the <a> tag with the source-url class—this is the link that appears as the original site on the archive page.
  • Fallback: If the primary method fails, it checks the og:url meta tag, which some archive snapshots use to define the original content URL.
  • Error Handling: Catches HTTP errors, timeouts, and cases where the original URL can't be found.

Approach 2: Use the Archive Redirect Endpoint (Less Stable)

Some archive.is variants support a redirect endpoint that sends you directly to the original URL. For example, adding /redirect/ before the snapshot hash: http://archive.is/redirect/9mIro. However, this endpoint isn't consistently supported across all archive domains, so it's less reliable than parsing the page.

Here's a quick example of this method:

import requests

def get_original_via_redirect(archive_url):
    redirect_url = archive_url.replace("archive.is/", "archive.is/redirect/")
    try:
        # Use HEAD request to avoid downloading the entire page
        response = requests.head(redirect_url, allow_redirects=True, timeout=10)
        return response.url
    except requests.exceptions.RequestException as e:
        raise RuntimeError(f"Redirect failed: {str(e)}") from e

Note: This method can break if archive sites change their redirect structure, so stick with Approach 1 for long-term reliability.


内容的提问来源于stack exchange,提问作者Joshua Meyers

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 07:58:47