如何使用BeautifulSoup提取网页中的非本地链接?
Got it, let's fix up that code to pull external links instead of internal ones. Here's a robust, reusable way to do it:
Extract External Links with BeautifulSoup
Instead of hardcoding a domain regex to target internal links, we'll dynamically parse the root URL's domain to filter out any links that belong to it. This makes the function work for any website you want to crawl.
Revised Code
import re from bs4 import BeautifulSoup from urllib import parse, request def get_external_links(root_url, html): # Parse the root URL to get its base domain (e.g., https://www.eecs.mitx.edu) root_parsed = parse.urlparse(root_url) root_domain = f"{root_parsed.scheme}://{root_parsed.netloc}" soup = BeautifulSoup(html, 'html.parser') external_links = [] # Loop through all <a> tags that have an href attribute for link in soup.find_all('a', href=True): href = link.get('href') # Parse the link to check its structure link_parsed = parse.urlparse(href) # Skip relative URLs (they're local to the domain) if not link_parsed.scheme and not link_parsed.netloc: continue # Resolve any partial links to their full absolute form full_link = parse.urljoin(root_url, href) full_parsed = parse.urlparse(full_link) full_domain = f"{full_parsed.scheme}://{full_parsed.netloc}" # If the link's domain doesn't match the root, it's external if full_domain != root_domain: # Grab the link text (fall back to the URL if there's no text) link_text = link.get_text(strip=True) or href external_links.append((full_link, link_text)) return external_links
Key Details Explained
- Dynamic Domain Handling: Using
urlparselets us avoid hardcoding the target domain—this function works for any root URL you pass in. - Filter Logic:
- Skip relative paths (like
/aboutor../contact) since they point to the same domain. - For absolute URLs, compare their full domain to the root domain. If they don't match, we add them to our external links list.
urljoinensures we resolve any tricky partial links (like//example.com) to their full absolute form.
- Skip relative paths (like
- Link Text Extraction: We pull the visible text of the link (cleaned of extra whitespace) for readability, using the URL itself as a fallback if there's no text.
Quick Usage Example
If you're fetching the HTML directly from a URL, here's how to use the function:
root_url = "https://www.eecs.mitx.edu/" response = request.urlopen(root_url) html = response.read().decode('utf-8') external_links = get_external_links(root_url, html) # Print out the results for link, text in external_links: print(f"Link Text: {text} | URL: {link}")
内容的提问来源于stack exchange,提问作者Tashi
相关产品推荐
相关产品推荐

