Python爬虫求助:提取Href链接并实现逐个链接爬取
Hey there! Great start on your web scraper—you’ve already nailed the first critical step: extracting all those restaurant links from the main page. Let’s break down exactly how to build on this and crawl each individual restaurant page to pull the data you need.
Step 1: Set Up Your CSV Properly
First, you’ll want to define clear headers for your CSV so your data is structured. Add this right after initializing your CSV writer:
# Write CSV headers f.writerow(['Restaurant Name', 'Address', 'Cuisine', 'Full Menu URL'])
Step 2: Loop Through Each Link & Crawl Details
Instead of just printing each completeurl, we’ll send a request to each page, parse its content, extract relevant details, and write them to your CSV. We’ll also add basic error handling and delays to avoid getting blocked.
Here’s how to modify your loop:
import time # Don't forget to import this at the top! for url2 in soup.find_all('h3', class_='restaurant__title'): completeurl = url1 + url2.a.get('href') print(f"Scraping: {completeurl}") try: # Send request to the restaurant page restaurant_response = requests.get(completeurl) # Add a delay to be respectful to the server time.sleep(2) # Parse the restaurant page restaurant_soup = BeautifulSoup(restaurant_response.text, "html.parser") # Extract data (adjust selectors based on what you need!) # Example selectors for menupages.com: name = restaurant_soup.find('h1', class_='restaurant-name').get_text(strip=True) if restaurant_soup.find('h1', class_='restaurant-name') else 'N/A' address = restaurant_soup.find('div', class_='restaurant-address').get_text(strip=True) if restaurant_soup.find('div', class_='restaurant-address') else 'N/A' cuisine = restaurant_soup.find('span', class_='restaurant-cuisine').get_text(strip=True) if restaurant_soup.find('span', class_='restaurant-cuisine') else 'N/A' # Write the row to CSV f.writerow([name, address, cuisine, completeurl]) except Exception as e: print(f"Error scraping {completeurl}: {str(e)}") continue
Step 3: Key Tips for Smooth Scraping
- Respect the Website’s Rules: Check the site’s
robots.txt(e.g.,https://menupages.com/robots.txt) to make sure scraping is allowed for the pages you’re targeting. - User-Agent Spoofing: Some sites block default
requestsuser agents. Add a custom header to your requests:headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} # Use this in your requests: response = requests.get(url, headers=headers) restaurant_response = requests.get(completeurl, headers=headers) - Handle Dynamic Content: If some data loads with JavaScript, you might need tools like
Seleniuminstead ofrequests—but start with the basics first, since menupages likely serves static HTML for restaurant details. - Close Files Properly: It’s better to use a
withstatement for your CSV file to ensure it closes correctly:with open('Restuarants_details.csv', 'w', newline='', encoding='utf-8') as csvfile: f = csv.writer(csvfile) # Rest of your code here (headers, loops, etc.)
Full Modified Code
Putting it all together, here’s your complete scraper:
import requests import re from bs4 import BeautifulSoup import csv import time url = 'https://menupages.com/restaurants/ny-new-york' url1 = 'https://menupages.com' headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'} response = requests.get(url, headers=headers) # Use with statement to manage CSV file with open('Restuarants_details.csv', 'w', newline='', encoding='utf-8') as csvfile: f = csv.writer(csvfile) # Write headers f.writerow(['Restaurant Name', 'Address', 'Cuisine', 'Full Menu URL']) soup = BeautifulSoup(response.text, "html.parser") for url2 in soup.find_all('h3', class_='restaurant__title'): completeurl = url1 + url2.a.get('href') print(f"Scraping: {completeurl}") try: restaurant_response = requests.get(completeurl, headers=headers) time.sleep(2) restaurant_soup = BeautifulSoup(restaurant_response.text, "html.parser") # Extract data name = restaurant_soup.find('h1', class_='restaurant-name').get_text(strip=True) if restaurant_soup.find('h1', class_='restaurant-name') else 'N/A' address = restaurant_soup.find('div', class_='restaurant-address').get_text(strip=True) if restaurant_soup.find('div', class_='restaurant-address') else 'N/A' cuisine = restaurant_soup.find('span', class_='restaurant-cuisine').get_text(strip=True) if restaurant_soup.find('span', class_='restaurant-cuisine') else 'N/A' f.writerow([name, address, cuisine, completeurl]) except Exception as e: print(f"Error scraping {completeurl}: {str(e)}") continue print("Scraping complete! Check your CSV file.")
内容的提问来源于stack exchange,提问作者Waleed Ahmed

