You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫求助:提取Href链接并实现逐个链接爬取

How to Extend Your Python Scraper to Crawl Restaurant Details

Hey there! Great start on your web scraper—you’ve already nailed the first critical step: extracting all those restaurant links from the main page. Let’s break down exactly how to build on this and crawl each individual restaurant page to pull the data you need.

Step 1: Set Up Your CSV Properly

First, you’ll want to define clear headers for your CSV so your data is structured. Add this right after initializing your CSV writer:

# Write CSV headers
f.writerow(['Restaurant Name', 'Address', 'Cuisine', 'Full Menu URL'])

Instead of just printing each completeurl, we’ll send a request to each page, parse its content, extract relevant details, and write them to your CSV. We’ll also add basic error handling and delays to avoid getting blocked.

Here’s how to modify your loop:

import time  # Don't forget to import this at the top!

for url2 in soup.find_all('h3', class_='restaurant__title'):
    completeurl = url1 + url2.a.get('href')
    print(f"Scraping: {completeurl}")
    
    try:
        # Send request to the restaurant page
        restaurant_response = requests.get(completeurl)
        # Add a delay to be respectful to the server
        time.sleep(2)
        
        # Parse the restaurant page
        restaurant_soup = BeautifulSoup(restaurant_response.text, "html.parser")
        
        # Extract data (adjust selectors based on what you need!)
        # Example selectors for menupages.com:
        name = restaurant_soup.find('h1', class_='restaurant-name').get_text(strip=True) if restaurant_soup.find('h1', class_='restaurant-name') else 'N/A'
        address = restaurant_soup.find('div', class_='restaurant-address').get_text(strip=True) if restaurant_soup.find('div', class_='restaurant-address') else 'N/A'
        cuisine = restaurant_soup.find('span', class_='restaurant-cuisine').get_text(strip=True) if restaurant_soup.find('span', class_='restaurant-cuisine') else 'N/A'
        
        # Write the row to CSV
        f.writerow([name, address, cuisine, completeurl])
        
    except Exception as e:
        print(f"Error scraping {completeurl}: {str(e)}")
        continue

Step 3: Key Tips for Smooth Scraping

  • Respect the Website’s Rules: Check the site’s robots.txt (e.g., https://menupages.com/robots.txt) to make sure scraping is allowed for the pages you’re targeting.
  • User-Agent Spoofing: Some sites block default requests user agents. Add a custom header to your requests:
    headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}
    # Use this in your requests:
    response = requests.get(url, headers=headers)
    restaurant_response = requests.get(completeurl, headers=headers)
    
  • Handle Dynamic Content: If some data loads with JavaScript, you might need tools like Selenium instead of requests—but start with the basics first, since menupages likely serves static HTML for restaurant details.
  • Close Files Properly: It’s better to use a with statement for your CSV file to ensure it closes correctly:
    with open('Restuarants_details.csv', 'w', newline='', encoding='utf-8') as csvfile:
        f = csv.writer(csvfile)
        # Rest of your code here (headers, loops, etc.)
    

Full Modified Code

Putting it all together, here’s your complete scraper:

import requests
import re
from bs4 import BeautifulSoup
import csv
import time

url = 'https://menupages.com/restaurants/ny-new-york'
url1 = 'https://menupages.com'
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/91.0.4472.124 Safari/537.36'}

response = requests.get(url, headers=headers)

# Use with statement to manage CSV file
with open('Restuarants_details.csv', 'w', newline='', encoding='utf-8') as csvfile:
    f = csv.writer(csvfile)
    # Write headers
    f.writerow(['Restaurant Name', 'Address', 'Cuisine', 'Full Menu URL'])
    
    soup = BeautifulSoup(response.text, "html.parser")
    
    for url2 in soup.find_all('h3', class_='restaurant__title'):
        completeurl = url1 + url2.a.get('href')
        print(f"Scraping: {completeurl}")
        
        try:
            restaurant_response = requests.get(completeurl, headers=headers)
            time.sleep(2)
            
            restaurant_soup = BeautifulSoup(restaurant_response.text, "html.parser")
            
            # Extract data
            name = restaurant_soup.find('h1', class_='restaurant-name').get_text(strip=True) if restaurant_soup.find('h1', class_='restaurant-name') else 'N/A'
            address = restaurant_soup.find('div', class_='restaurant-address').get_text(strip=True) if restaurant_soup.find('div', class_='restaurant-address') else 'N/A'
            cuisine = restaurant_soup.find('span', class_='restaurant-cuisine').get_text(strip=True) if restaurant_soup.find('span', class_='restaurant-cuisine') else 'N/A'
            
            f.writerow([name, address, cuisine, completeurl])
            
        except Exception as e:
            print(f"Error scraping {completeurl}: {str(e)}")
            continue

print("Scraping complete! Check your CSV file.")

内容的提问来源于stack exchange,提问作者Waleed Ahmed

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 08:10:10