You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Beautiful Soup避免爬取链接重复:检查queue与crawled文件

Hey there! It looks like your crawler is struggling with duplicate links because it only removes duplicates within a single page, not checking against links already stored in queue.txt or crawled.txt. Let's refactor your code to fix this for good, with a focus on efficient link tracking.

Core Problem Breakdown

  • Your current approach only deduplicates links from one page before writing, but doesn't compare against links that are already in your queue or already crawled.
  • Using lists for link storage makes duplicate checks slow; sets are far better here since they have near-instant membership checks and automatically enforce uniqueness.

Solution Strategy

  1. Load existing state: Pull all links from queue.txt and crawled.txt into sets at the start of each crawl cycle.
  2. Process queue links: Take a link from the queue, mark it as crawled, then extract new links from its page.
  3. Filter new links: Only add a link to the queue if it's not already in the queue set or crawled set.
  4. Persist state: Update the queue and crawled files after each crawl to maintain progress between runs.

Refactored Code

from bs4 import BeautifulSoup
import requests
from urllib.parse import urlparse

def load_links_from_file(file_path):
    """Load links from a file into a set (auto-removes duplicates)"""
    links = set()
    try:
        with open(file_path, 'r') as f:
            for line in f:
                link = line.strip()
                if link:  # Skip empty lines
                    links.add(link)
    except FileNotFoundError:
        # Return empty set if file doesn't exist yet
        pass
    return links

def save_links_to_file(links, file_path):
    """Save a set of links to a file, overwriting existing content"""
    with open(file_path, 'w') as f:
        for link in links:
            f.write(f"{link}\n")

def normalize_link(base_url, link):
    """Convert relative/partial links to full valid URLs"""
    if not link:
        return None
    
    link = str(link).strip()
    # Skip anchor links, javascript, mailto, and other non-web links
    skip_prefixes = ('#', 'javascript:', 'mailto:', 'tel:', 'ftp:')
    if link.startswith(skip_prefixes):
        return None
    
    parsed_base = urlparse(base_url)
    domain = f"{parsed_base.scheme}://{parsed_base.netloc}"
    
    # Handle protocol-relative links (//example.com)
    if link.startswith('//'):
        return f"https:{link}"
    # Handle root-relative links (/path/to/page)
    elif link.startswith('/'):
        return f"{domain}{link}"
    # Handle full HTTP/HTTPS links
    elif link.startswith(('http://', 'https://')):
        return link
    # Ignore other invalid formats
    else:
        return None

def get_valid_links(base_url):
    """Extract and normalize all valid links from a page"""
    try:
        response = requests.get(base_url, timeout=10)
        response.raise_for_status()  # Raise error for HTTP 4xx/5xx responses
        soup = BeautifulSoup(response.content, 'html.parser')
        
        valid_links = set()
        # Only target <a> tags with an href attribute
        for a_tag in soup.find_all('a', href=True):
            raw_link = a_tag.get('href')
            normalized_link = normalize_link(base_url, raw_link)
            if normalized_link:
                valid_links.add(normalized_link)
        
        return valid_links
    except Exception as e:
        print(f"Failed to crawl {base_url}: {str(e)}")
        return set()

def run_crawler(start_url):
    # Initialize queue and crawled sets with existing links
    queue = load_links_from_file('queue.txt')
    crawled = load_links_from_file('crawled.txt')
    
    # Add start URL if it's not already tracked
    if start_url not in queue and start_url not in crawled:
        queue.add(start_url)
    
    while queue:
        # Grab the next link to crawl (pop removes it from the queue)
        current_link = queue.pop()
        if current_link in crawled:
            continue
        
        print(f"Crawling: {current_link}")
        
        # Mark the link as crawled
        crawled.add(current_link)
        
        # Get all valid links from the current page
        new_links = get_valid_links(current_link)
        
        # Add only new, unique links to the queue
        for link in new_links:
            if link not in queue and link not in crawled:
                queue.add(link)
                print(f"Added to queue: {link}")
        
        # Update files to save progress
        save_links_to_file(queue, 'queue.txt')
        save_links_to_file(crawled, 'crawled.txt')
    
    print("Crawl completed!")

# Start the crawler with your target URL
run_crawler('https://stackabuse.com/reading-and-writing-lists-to-a-file-in-python/')

Key Improvements Explained

  • Set-based tracking: Using sets for queue and crawled ensures no duplicates and makes checking existing links lightning fast.
  • Centralized link normalization: The normalize_link function cleans up messy relative/protocol-relative links in one place, replacing your scattered if checks.
  • Persistent state: load_links_from_file and save_links_to_file let the crawler resume where it left off if you restart it.
  • Robust error handling: Added timeouts and HTTP error checks to avoid crashes from unresponsive pages.
  • Cleaner workflow: The run_crawler function organizes the crawl into a clear loop, making it easy to follow and modify.

Quick Extra Tips

  • Add a time.sleep(1) (import time first) between requests to avoid overwhelming the target server and getting blocked.
  • If you want to limit crawl depth (how many levels from the start URL to go), add a depth counter and stop when it reaches your limit.

内容的提问来源于stack exchange,提问作者Koomcravet

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.09 20:57:35