如何用Python 3+Beautiful Soup批量抓取多URL及读取CSV中的URL
Got it, let's break down how to adapt your existing code to handle batch URL scraping from a CSV file. We'll go step by step to make this straightforward.
Step 1: Set Up Your CSV File
First, create a CSV file (let's name it urls_to_scrape.csv) with one URL per row. A simple, clean structure would look like this:
url https://www.tennis-point.fr/index.php?stoken=737F2976&lang=1&cl=search&searchparam=E705Y-0193 https://www.tennis-point.fr/index.php?stoken=737F2976&lang=1&cl=search&searchparam=ANOTHER_PRODUCT_CODE https://www.tennis-point.fr/index.php?stoken=737F2976&lang=1&cl=search&searchparam=YOUR_THIRD_PRODUCT
Double-check that each URL has no extra spaces or formatting issues.
Step 2: Refactor Your Scraper into a Reusable Function
Your existing code works perfectly for single URLs—let's wrap that logic into a function so we can call it repeatedly for every URL in the CSV. We'll also add basic error handling to avoid crashes if a URL fails to load or the page structure changes.
from bs4 import BeautifulSoup import requests import csv def extract_product_links(search_url): """Extracts product href links from a tennis-point search results page""" try: # Fetch the webpage response = requests.get(search_url) response.raise_for_status() # Trigger an error for bad HTTP codes (404, 500, etc.) # Parse the page with BeautifulSoup soup = BeautifulSoup(response.text, 'lxml') # Locate the products container and extract links products_container = soup.find("div", {"class": "productsPicture"}) if not products_container: print(f"No products found for URL: {search_url}") return [] tags = products_container.findAll("a") links = [tag.get('href') for tag in tags] return links except requests.exceptions.RequestException as e: print(f"Error fetching {search_url}: {str(e)}") return []
Step 3: Read CSV & Process URLs in Batch
Now we'll write code to read the CSV file, loop through each URL, and run our scraper function. We'll also add optional code to save the results to a new CSV so you can keep track of all extracted links.
def main(): # Path to your input CSV with search URLs input_csv = "urls_to_scrape.csv" # Path to save extracted links (optional but useful) output_csv = "extracted_links.csv" # Open input CSV and read URLs with open(input_csv, mode='r', newline='', encoding='utf-8') as infile: reader = csv.DictReader(infile) # Prepare to write results to output CSV with open(output_csv, mode='w', newline='', encoding='utf-8') as outfile: writer = csv.writer(outfile) writer.writerow(["Search URL", "Extracted Link"]) # Header row for row in reader: search_url = row['url'].strip() if not search_url: continue # Skip empty rows print(f"Processing URL: {search_url}") extracted_links = extract_product_links(search_url) # Write each extracted link to the output CSV for link in extracted_links: writer.writerow([search_url, link]) print(f"Extracted {len(extracted_links)} links for this URL\n") if __name__ == "__main__": main()
Extra Tips for Smooth Scraping
- Add delays: To avoid overwhelming the server and getting blocked, add a small delay between requests with
time.sleep(2)(don't forget to import thetimemodule first). - Custom User-Agent: Some sites block the default
requestsuser agent. Add a browser-like header to yourrequests.getcall:headers = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"} response = requests.get(search_url, headers=headers) - Handle dynamic content: If the site uses JavaScript to load products later, you might need tools like
Seleniuminstead ofrequests—but based on your original code, static content scraping should work fine here.
内容的提问来源于stack exchange,提问作者AnotherUser31

