You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python 3+Beautiful Soup批量抓取多URL及读取CSV中的URL

Got it, let's break down how to adapt your existing code to handle batch URL scraping from a CSV file. We'll go step by step to make this straightforward.

Step 1: Set Up Your CSV File

First, create a CSV file (let's name it urls_to_scrape.csv) with one URL per row. A simple, clean structure would look like this:

url
https://www.tennis-point.fr/index.php?stoken=737F2976&lang=1&cl=search&searchparam=E705Y-0193
https://www.tennis-point.fr/index.php?stoken=737F2976&lang=1&cl=search&searchparam=ANOTHER_PRODUCT_CODE
https://www.tennis-point.fr/index.php?stoken=737F2976&lang=1&cl=search&searchparam=YOUR_THIRD_PRODUCT

Double-check that each URL has no extra spaces or formatting issues.

Step 2: Refactor Your Scraper into a Reusable Function

Your existing code works perfectly for single URLs—let's wrap that logic into a function so we can call it repeatedly for every URL in the CSV. We'll also add basic error handling to avoid crashes if a URL fails to load or the page structure changes.

from bs4 import BeautifulSoup
import requests
import csv

def extract_product_links(search_url):
    """Extracts product href links from a tennis-point search results page"""
    try:
        # Fetch the webpage
        response = requests.get(search_url)
        response.raise_for_status()  # Trigger an error for bad HTTP codes (404, 500, etc.)
        
        # Parse the page with BeautifulSoup
        soup = BeautifulSoup(response.text, 'lxml')
        
        # Locate the products container and extract links
        products_container = soup.find("div", {"class": "productsPicture"})
        if not products_container:
            print(f"No products found for URL: {search_url}")
            return []
        
        tags = products_container.findAll("a")
        links = [tag.get('href') for tag in tags]
        return links
    
    except requests.exceptions.RequestException as e:
        print(f"Error fetching {search_url}: {str(e)}")
        return []

Step 3: Read CSV & Process URLs in Batch

Now we'll write code to read the CSV file, loop through each URL, and run our scraper function. We'll also add optional code to save the results to a new CSV so you can keep track of all extracted links.

def main():
    # Path to your input CSV with search URLs
    input_csv = "urls_to_scrape.csv"
    # Path to save extracted links (optional but useful)
    output_csv = "extracted_links.csv"
    
    # Open input CSV and read URLs
    with open(input_csv, mode='r', newline='', encoding='utf-8') as infile:
        reader = csv.DictReader(infile)
        
        # Prepare to write results to output CSV
        with open(output_csv, mode='w', newline='', encoding='utf-8') as outfile:
            writer = csv.writer(outfile)
            writer.writerow(["Search URL", "Extracted Link"])  # Header row
            
            for row in reader:
                search_url = row['url'].strip()
                if not search_url:
                    continue  # Skip empty rows
                
                print(f"Processing URL: {search_url}")
                extracted_links = extract_product_links(search_url)
                
                # Write each extracted link to the output CSV
                for link in extracted_links:
                    writer.writerow([search_url, link])
                
                print(f"Extracted {len(extracted_links)} links for this URL\n")

if __name__ == "__main__":
    main()

Extra Tips for Smooth Scraping

  • Add delays: To avoid overwhelming the server and getting blocked, add a small delay between requests with time.sleep(2) (don't forget to import the time module first).
  • Custom User-Agent: Some sites block the default requests user agent. Add a browser-like header to your requests.get call:
    headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"}
    response = requests.get(search_url, headers=headers)
    
  • Handle dynamic content: If the site uses JavaScript to load products later, you might need tools like Selenium instead of requests—but based on your original code, static content scraping should work fine here.

内容的提问来源于stack exchange,提问作者AnotherUser31

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 06:38:20