You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何基于现有URL解析函数实现同域名内链的递归遍历?

嘿,这本质就是要做一个限定域名的深度爬虫嘛,最简单的路子就是结合「链接去重」+「深度/广度遍历」,我给你捋捋最直白的实现方式:

核心思路拆解

首先得把几个关键逻辑拎清楚,不然爬着爬着就乱套了:

  • 锁死目标域名:不管解析出的是相对路径还是绝对链接,都得先判断是不是和初始域名同属一个根域,坚决过滤外部链接。
  • 强制去重:用集合存已经爬过的链接——同一个链接可能在N个页面出现,重复处理纯浪费时间。
  • 遍历方式二选一:递归写法最简洁,但深层级网站容易栈溢出;迭代用队列做广度优先,稳得一批,适合复杂站点。
具体实现示例(Python)

假设你的解析函数还没完全处理相对链接这些细节,我把完整的逻辑写出来,你可以直接套:

第一步:完善内部链接解析函数

这个函数负责从指定页面抠出所有同域名的内部链接,还会自动处理相对路径转绝对路径、标准化URL去重:

from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup

def get_internal_links(base_url, page_url):
    """解析页面,返回去重后的同域名内部链接列表"""
    # 先处理页面访问异常,避免爬虫直接崩
    try:
        response = requests.get(page_url, timeout=10)
        response.raise_for_status()  # 捕获4xx/5xx状态码
    except Exception as e:
        print(f"跳过无法访问的页面: {page_url} | 错误: {e}")
        return []
    
    # 解析页面所有<a>标签的链接
    soup = BeautifulSoup(response.text, "html.parser")
    all_a_tags = soup.find_all("a", href=True)
    
    base_domain = urlparse(base_url).netloc  # 提取目标域名,比如example.com
    internal_links = set()  # 用集合自动去重
    
    for tag in all_a_tags:
        raw_href = tag["href"]
        # 把相对链接转成绝对链接(比如"/about"转成"https://example.com/about")
        absolute_url = urljoin(page_url, raw_href)
        # 提取当前链接的域名,判断是否和目标域名一致
        link_domain = urlparse(absolute_url).netloc
        
        if link_domain == base_domain:
            # 标准化URL:去掉末尾斜杠,避免example.com/page和example.com/page/被当成两个链接
            normalized_url = absolute_url.rstrip("/")
            internal_links.add(normalized_url)
    
    return list(internal_links)

第二步:爬虫主逻辑(两种写法)

写法1:递归实现(代码最简洁,适合层级不深的站点)

# 用全局集合存已爬链接,避免重复处理
crawled_links = set()

def crawl_recursive(base_url, current_url):
    # 如果已经爬过,直接跳过
    if current_url in crawled_links:
        return
    
    print(f"正在爬取: {current_url}")
    crawled_links.add(current_url)
    
    # 获取当前页面的所有内部链接
    new_links = get_internal_links(base_url, current_url)
    
    # 递归处理每个新链接
    for link in new_links:
        crawl_recursive(base_url, link)

# 启动爬虫,比如爬example.com的全站
crawl_recursive("https://example.com", "https://example.com")

写法2:迭代实现(用队列做广度优先,适合深层级大站点)

如果目标网站层级特别深(比如有几十层嵌套链接),递归会触发RecursionError,这时候用队列的迭代写法更稳妥:

from collections import deque

def crawl_iterative(base_url, start_url):
    crawled_links = set()
    # 用队列存待爬的链接,先进先出就是广度优先遍历
    pending_links = deque([start_url])
    
    while pending_links:
        current_url = pending_links.popleft()
        if current_url in crawled_links:
            continue
        
        print(f"正在爬取: {current_url}")
        crawled_links.add(current_url)
        
        # 获取当前页面的内部链接
        new_links = get_internal_links(base_url, current_url)
        
        # 把未爬过的链接加入队列
        for link in new_links:
            if link not in crawled_links and link not in pending_links:
                pending_links.append(link)

# 启动爬虫
crawl_iterative("https://example.com", "https://example.com")
必看的踩坑点
  • URL标准化一定要做:比如去掉末尾斜杠、统一协议(https/http),不然会出现大量重复链接。
  • 加速率限制:爬别人的网站时,记得加time.sleep(1)之类的延迟,不然很容易被IP封禁。
  • 遵守robots.txt:先看看目标网站的robots.txt(比如https://example.com/robots.txt),别爬人家明确禁止的路径。
  • 异常处理要到位:网页可能404、超时、反爬,一定要捕获异常,不然爬虫半路就挂了。

内容的提问来源于stack exchange,提问作者user7496746

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 09:01:55