如何编写Python代码递归提取网页内部href子链接?
Recursively Extract Internal Links from 3gpp.org
Your current code does a great job pulling all anchor links from the homepage, but to focus on internal links and crawl them recursively, we need to add a few key features: filtering internal URLs, tracking visited pages to avoid loops, and handling relative paths properly. Here's a revised implementation:
Improved Code
from urllib.request import urlopen, Request from urllib.parse import urljoin from bs4 import BeautifulSoup import time # Base URL of the target site BASE_URL = "https://www.3gpp.org/" # Set to track visited URLs and prevent infinite loops visited_urls = set() def crawl_internal_links(url): # Skip if we've already visited this URL if url in visited_urls: return print(f"Crawling: {url}") visited_urls.add(url) try: # Add a user-agent header to avoid being blocked by the server request = Request(url, headers={"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36"}) response = urlopen(request) # Only process HTML content content_type = response.getheader("Content-Type") if not content_type or "text/html" not in content_type: return soup = BeautifulSoup(response, "lxml") # Iterate through all anchor tags for anchor in soup.find_all("a"): href = anchor.get("href") # Skip empty or non-HTTP links (like mailto, tel, anchors) if not href or href.startswith(("mailto:", "tel:", "#", "javascript:")): continue # Convert relative URLs to absolute paths absolute_url = urljoin(BASE_URL, href) # Check if the URL is internal (belongs to 3gpp.org) if absolute_url.startswith(BASE_URL): # Add a small delay to be polite to the server time.sleep(1) crawl_internal_links(absolute_url) except Exception as e: print(f"Failed to crawl {url}: {str(e)}") # Start crawling from the base URL crawl_internal_links(BASE_URL)
Key Improvements Explained
- Internal Link Filtering: We use
urljointo convert relative paths (like/about/) to full absolute URLs, then check if they start with our base URL to confirm they're internal. - Loop Prevention: The
visited_urlsset keeps track of every page we've crawled, so we don't revisit the same URL infinitely. - Polite Crawling: Adding
time.sleep(1)between requests ensures we don't overload the server, and a user-agent header helps avoid being blocked as a bot. - Error Handling: Basic exception catching ensures the script doesn't crash if it encounters a broken link or server error.
- Content Type Check: We skip non-HTML files (like PDFs or images) since they don't contain links to crawl.
Notes to Consider
- For very large sites, recursive calls might hit Python's recursion depth limit. In that case, you could switch to an iterative approach using a queue (like
collections.deque) instead. - You might want to expand the filter to exclude certain paths (like administrative pages) by checking if the URL contains specific strings.
- Always respect the site's
robots.txtfile (check it athttps://www.3gpp.org/robots.txt) to make sure you're allowed to crawl their pages.
内容的提问来源于stack exchange,提问作者Abdul Raoof
相关产品推荐
相关产品推荐

