You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用BeautifulSoup提取网页中的非本地链接?

Got it, let's fix up that code to pull external links instead of internal ones. Here's a robust, reusable way to do it:

Instead of hardcoding a domain regex to target internal links, we'll dynamically parse the root URL's domain to filter out any links that belong to it. This makes the function work for any website you want to crawl.

Revised Code

import re
from bs4 import BeautifulSoup
from urllib import parse, request

def get_external_links(root_url, html):
    # Parse the root URL to get its base domain (e.g., https://www.eecs.mitx.edu)
    root_parsed = parse.urlparse(root_url)
    root_domain = f"{root_parsed.scheme}://{root_parsed.netloc}"
    
    soup = BeautifulSoup(html, 'html.parser')
    external_links = []
    
    # Loop through all <a> tags that have an href attribute
    for link in soup.find_all('a', href=True):
        href = link.get('href')
        # Parse the link to check its structure
        link_parsed = parse.urlparse(href)
        
        # Skip relative URLs (they're local to the domain)
        if not link_parsed.scheme and not link_parsed.netloc:
            continue
        
        # Resolve any partial links to their full absolute form
        full_link = parse.urljoin(root_url, href)
        full_parsed = parse.urlparse(full_link)
        full_domain = f"{full_parsed.scheme}://{full_parsed.netloc}"
        
        # If the link's domain doesn't match the root, it's external
        if full_domain != root_domain:
            # Grab the link text (fall back to the URL if there's no text)
            link_text = link.get_text(strip=True) or href
            external_links.append((full_link, link_text))
    
    return external_links

Key Details Explained

  • Dynamic Domain Handling: Using urlparse lets us avoid hardcoding the target domain—this function works for any root URL you pass in.
  • Filter Logic:
    • Skip relative paths (like /about or ../contact) since they point to the same domain.
    • For absolute URLs, compare their full domain to the root domain. If they don't match, we add them to our external links list.
    • urljoin ensures we resolve any tricky partial links (like //example.com) to their full absolute form.
  • Link Text Extraction: We pull the visible text of the link (cleaned of extra whitespace) for readability, using the URL itself as a fallback if there's no text.

Quick Usage Example

If you're fetching the HTML directly from a URL, here's how to use the function:

root_url = "https://www.eecs.mitx.edu/"
response = request.urlopen(root_url)
html = response.read().decode('utf-8')
external_links = get_external_links(root_url, html)

# Print out the results
for link, text in external_links:
    print(f"Link Text: {text} | URL: {link}")

内容的提问来源于stack exchange,提问作者Tashi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.25 04:27:15