如何调整Python3 CSV爬虫以支持多列提取及指定div类
Modified Python CSV Crawler for Multi-column Extraction with Targeted div Classes
I get it—you need to expand your existing CSV crawler to pull data from multiple columns, each paired with specific div classes. Let’s adjust the code to meet those requirements while keeping it flexible and robust.
First, Let’s Start with a Plausible Version of Your Original Code
This matches the behavior you described (extracting from row[0] with a single div class):
import csv import requests from bs4 import BeautifulSoup def crawl_url(url, div_tag_class): try: response = requests.get(url) soup = BeautifulSoup(response.text, 'html.parser') target_div = soup.find('div', class_=div_tag_class) return target_div.get_text(strip=True) if target_div else "No content found" except Exception as e: return f"Error: {str(e)}" def main(): with open('urls.csv', 'r', newline='', encoding='utf-8') as csvfile: reader = csv.reader(csvfile) next(reader) # Skip header for row in reader: url = row[0] content = crawl_url(url, "default_class") print(f"URL: {url} | Content: {content}") if __name__ == "__main__": main()
Modified Code for Multi-column & Class-specific Extraction
This version lets you define exactly which columns use which div classes, and handles multiple columns seamlessly:
import csv import requests from bs4 import BeautifulSoup def crawl_url(url, div_tag_class): if not url: # Skip empty cells to avoid unnecessary requests return "Empty URL" try: response = requests.get(url, timeout=10) response.raise_for_status() # Catch HTTP errors (404, 500, etc.) soup = BeautifulSoup(response.text, 'html.parser') target_div = soup.find('div', class_=div_tag_class) return target_div.get_text(strip=True) if target_div else "No matching div found" except requests.exceptions.RequestException as e: return f"Request failed: {str(e)}" except Exception as e: return f"Unexpected error: {str(e)}" def main(): # Define your column-to-div-class mapping here # Format: { 'div_class_name': [list_of_column_indices] } column_class_mapping = { "productsPicture": [1, 2], # Columns 1 and 2 use this class "product_content": [4, 5] # Columns 4 and 5 use this class } with open('urls.csv', 'r', newline='', encoding='utf-8') as csvfile: reader = csv.reader(csvfile) header = next(reader) # Capture header for more readable output for row_num, row in enumerate(reader, start=2): # Start counting rows after header print(f"\n=== Processing Row {row_num} ===") # Iterate over each class and its associated columns for div_class, columns in column_class_mapping.items(): for col_idx in columns: # Avoid index errors if the row is shorter than expected if col_idx >= len(row): print(f" Column {col_idx} (class: {div_class}): Out of bounds") continue url = row[col_idx] content = crawl_url(url, div_class) # Use header name if available, else "Unnamed Column" col_name = header[col_idx] if col_idx < len(header) else f"Column {col_idx}" print(f" {col_name} (class: {div_class}):") print(f" URL: {url}") # Truncate long content for readability print(f" Content: {content[:100]}..." if len(content) > 100 else f" Content: {content}") if __name__ == "__main__": main()
Key Improvements Explained
- Flexible Mapping: The
column_class_mappingdictionary lets you easily add/remove column-class pairs without rewriting core logic. Just update this dict to match your needs. - Robust Error Handling: Added checks for empty URLs, HTTP errors, and out-of-bounds columns to prevent crashes and give clear feedback.
- Readable Output: Includes row numbers, column names (from your CSV header), and truncated content to make results easier to scan.
- Scalable: Works with any number of columns and class pairs—no hardcoded limits.
Example urls.csv Input
ID,Product Image URL,Thumbnail URL,Product Description,Specs URL,Manual URL 1,https://example.com/product1/img,https://example.com/product1/thumb,"Wireless Headphones",https://example.com/product1/specs,https://example.com/product1/manual 2,https://example.com/product2/img,https://example.com/product2/thumb,"Bluetooth Speaker",https://example.com/product2/specs,https://example.com/product2/manual
Quick Setup Notes
- Install required packages if you haven’t already:
pip install requests beautifulsoup4 - Adjust the
timeoutvalue inrequests.get()if you need longer wait times for slow URLs. - If you want to save results to a new CSV instead of printing, add a
csv.writerobject and write rows with the collected content.
内容的提问来源于stack exchange,提问作者AnotherUser31
相关产品推荐
相关产品推荐

