如何用Beautiful Soup避免爬取链接重复:检查queue与crawled文件
Fixing Duplicate Link Issues in Your Beautiful Soup Web Crawler
Hey there! It looks like your crawler is struggling with duplicate links because it only removes duplicates within a single page, not checking against links already stored in queue.txt or crawled.txt. Let's refactor your code to fix this for good, with a focus on efficient link tracking.
Core Problem Breakdown
- Your current approach only deduplicates links from one page before writing, but doesn't compare against links that are already in your queue or already crawled.
- Using lists for link storage makes duplicate checks slow; sets are far better here since they have near-instant membership checks and automatically enforce uniqueness.
Solution Strategy
- Load existing state: Pull all links from
queue.txtandcrawled.txtinto sets at the start of each crawl cycle. - Process queue links: Take a link from the queue, mark it as crawled, then extract new links from its page.
- Filter new links: Only add a link to the queue if it's not already in the queue set or crawled set.
- Persist state: Update the queue and crawled files after each crawl to maintain progress between runs.
Refactored Code
from bs4 import BeautifulSoup import requests from urllib.parse import urlparse def load_links_from_file(file_path): """Load links from a file into a set (auto-removes duplicates)""" links = set() try: with open(file_path, 'r') as f: for line in f: link = line.strip() if link: # Skip empty lines links.add(link) except FileNotFoundError: # Return empty set if file doesn't exist yet pass return links def save_links_to_file(links, file_path): """Save a set of links to a file, overwriting existing content""" with open(file_path, 'w') as f: for link in links: f.write(f"{link}\n") def normalize_link(base_url, link): """Convert relative/partial links to full valid URLs""" if not link: return None link = str(link).strip() # Skip anchor links, javascript, mailto, and other non-web links skip_prefixes = ('#', 'javascript:', 'mailto:', 'tel:', 'ftp:') if link.startswith(skip_prefixes): return None parsed_base = urlparse(base_url) domain = f"{parsed_base.scheme}://{parsed_base.netloc}" # Handle protocol-relative links (//example.com) if link.startswith('//'): return f"https:{link}" # Handle root-relative links (/path/to/page) elif link.startswith('/'): return f"{domain}{link}" # Handle full HTTP/HTTPS links elif link.startswith(('http://', 'https://')): return link # Ignore other invalid formats else: return None def get_valid_links(base_url): """Extract and normalize all valid links from a page""" try: response = requests.get(base_url, timeout=10) response.raise_for_status() # Raise error for HTTP 4xx/5xx responses soup = BeautifulSoup(response.content, 'html.parser') valid_links = set() # Only target <a> tags with an href attribute for a_tag in soup.find_all('a', href=True): raw_link = a_tag.get('href') normalized_link = normalize_link(base_url, raw_link) if normalized_link: valid_links.add(normalized_link) return valid_links except Exception as e: print(f"Failed to crawl {base_url}: {str(e)}") return set() def run_crawler(start_url): # Initialize queue and crawled sets with existing links queue = load_links_from_file('queue.txt') crawled = load_links_from_file('crawled.txt') # Add start URL if it's not already tracked if start_url not in queue and start_url not in crawled: queue.add(start_url) while queue: # Grab the next link to crawl (pop removes it from the queue) current_link = queue.pop() if current_link in crawled: continue print(f"Crawling: {current_link}") # Mark the link as crawled crawled.add(current_link) # Get all valid links from the current page new_links = get_valid_links(current_link) # Add only new, unique links to the queue for link in new_links: if link not in queue and link not in crawled: queue.add(link) print(f"Added to queue: {link}") # Update files to save progress save_links_to_file(queue, 'queue.txt') save_links_to_file(crawled, 'crawled.txt') print("Crawl completed!") # Start the crawler with your target URL run_crawler('https://stackabuse.com/reading-and-writing-lists-to-a-file-in-python/')
Key Improvements Explained
- Set-based tracking: Using sets for
queueandcrawledensures no duplicates and makes checking existing links lightning fast. - Centralized link normalization: The
normalize_linkfunction cleans up messy relative/protocol-relative links in one place, replacing your scatteredifchecks. - Persistent state:
load_links_from_fileandsave_links_to_filelet the crawler resume where it left off if you restart it. - Robust error handling: Added timeouts and HTTP error checks to avoid crashes from unresponsive pages.
- Cleaner workflow: The
run_crawlerfunction organizes the crawl into a clear loop, making it easy to follow and modify.
Quick Extra Tips
- Add a
time.sleep(1)(importtimefirst) between requests to avoid overwhelming the target server and getting blocked. - If you want to limit crawl depth (how many levels from the start URL to go), add a depth counter and stop when it reaches your limit.
内容的提问来源于stack exchange,提问作者Koomcravet
相关产品推荐
相关产品推荐

