You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何编写Python代码递归提取网页内部href子链接?

Your current code does a great job pulling all anchor links from the homepage, but to focus on internal links and crawl them recursively, we need to add a few key features: filtering internal URLs, tracking visited pages to avoid loops, and handling relative paths properly. Here's a revised implementation:

Improved Code

from urllib.request import urlopen, Request
from urllib.parse import urljoin
from bs4 import BeautifulSoup
import time

# Base URL of the target site
BASE_URL = "https://www.3gpp.org/"
# Set to track visited URLs and prevent infinite loops
visited_urls = set()

def crawl_internal_links(url):
    # Skip if we've already visited this URL
    if url in visited_urls:
        return
    
    print(f"Crawling: {url}")
    visited_urls.add(url)
    
    try:
        # Add a user-agent header to avoid being blocked by the server
        request = Request(url, headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"})
        response = urlopen(request)
        
        # Only process HTML content
        content_type = response.getheader("Content-Type")
        if not content_type or "text/html" not in content_type:
            return
        
        soup = BeautifulSoup(response, "lxml")
        
        # Iterate through all anchor tags
        for anchor in soup.find_all("a"):
            href = anchor.get("href")
            
            # Skip empty or non-HTTP links (like mailto, tel, anchors)
            if not href or href.startswith(("mailto:", "tel:", "#", "javascript:")):
                continue
            
            # Convert relative URLs to absolute paths
            absolute_url = urljoin(BASE_URL, href)
            
            # Check if the URL is internal (belongs to 3gpp.org)
            if absolute_url.startswith(BASE_URL):
                # Add a small delay to be polite to the server
                time.sleep(1)
                crawl_internal_links(absolute_url)
                
    except Exception as e:
        print(f"Failed to crawl {url}: {str(e)}")

# Start crawling from the base URL
crawl_internal_links(BASE_URL)

Key Improvements Explained

  • Internal Link Filtering: We use urljoin to convert relative paths (like /about/) to full absolute URLs, then check if they start with our base URL to confirm they're internal.
  • Loop Prevention: The visited_urls set keeps track of every page we've crawled, so we don't revisit the same URL infinitely.
  • Polite Crawling: Adding time.sleep(1) between requests ensures we don't overload the server, and a user-agent header helps avoid being blocked as a bot.
  • Error Handling: Basic exception catching ensures the script doesn't crash if it encounters a broken link or server error.
  • Content Type Check: We skip non-HTML files (like PDFs or images) since they don't contain links to crawl.

Notes to Consider

  • For very large sites, recursive calls might hit Python's recursion depth limit. In that case, you could switch to an iterative approach using a queue (like collections.deque) instead.
  • You might want to expand the filter to exclude certain paths (like administrative pages) by checking if the URL contains specific strings.
  • Always respect the site's robots.txt file (check it at https://www.3gpp.org/robots.txt) to make sure you're allowed to crawl their pages.

内容的提问来源于stack exchange,提问作者Abdul Raoof

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.12 05:12:22