You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何调整Python3 CSV爬虫以支持多列提取及指定div类

Modified Python CSV Crawler for Multi-column Extraction with Targeted div Classes

I get it—you need to expand your existing CSV crawler to pull data from multiple columns, each paired with specific div classes. Let’s adjust the code to meet those requirements while keeping it flexible and robust.

First, Let’s Start with a Plausible Version of Your Original Code

This matches the behavior you described (extracting from row[0] with a single div class):

import csv
import requests
from bs4 import BeautifulSoup

def crawl_url(url, div_tag_class):
    try:
        response = requests.get(url)
        soup = BeautifulSoup(response.text, 'html.parser')
        target_div = soup.find('div', class_=div_tag_class)
        return target_div.get_text(strip=True) if target_div else "No content found"
    except Exception as e:
        return f"Error: {str(e)}"

def main():
    with open('urls.csv', 'r', newline='', encoding='utf-8') as csvfile:
        reader = csv.reader(csvfile)
        next(reader)  # Skip header
        for row in reader:
            url = row[0]
            content = crawl_url(url, "default_class")
            print(f"URL: {url} | Content: {content}")

if __name__ == "__main__":
    main()

Modified Code for Multi-column & Class-specific Extraction

This version lets you define exactly which columns use which div classes, and handles multiple columns seamlessly:

import csv
import requests
from bs4 import BeautifulSoup

def crawl_url(url, div_tag_class):
    if not url:  # Skip empty cells to avoid unnecessary requests
        return "Empty URL"
    try:
        response = requests.get(url, timeout=10)
        response.raise_for_status()  # Catch HTTP errors (404, 500, etc.)
        soup = BeautifulSoup(response.text, 'html.parser')
        target_div = soup.find('div', class_=div_tag_class)
        return target_div.get_text(strip=True) if target_div else "No matching div found"
    except requests.exceptions.RequestException as e:
        return f"Request failed: {str(e)}"
    except Exception as e:
        return f"Unexpected error: {str(e)}"

def main():
    # Define your column-to-div-class mapping here
    # Format: { 'div_class_name': [list_of_column_indices] }
    column_class_mapping = {
        "productsPicture": [1, 2],   # Columns 1 and 2 use this class
        "product_content": [4, 5]    # Columns 4 and 5 use this class
    }

    with open('urls.csv', 'r', newline='', encoding='utf-8') as csvfile:
        reader = csv.reader(csvfile)
        header = next(reader)  # Capture header for more readable output
        
        for row_num, row in enumerate(reader, start=2):  # Start counting rows after header
            print(f"\n=== Processing Row {row_num} ===")
            # Iterate over each class and its associated columns
            for div_class, columns in column_class_mapping.items():
                for col_idx in columns:
                    # Avoid index errors if the row is shorter than expected
                    if col_idx >= len(row):
                        print(f"  Column {col_idx} (class: {div_class}): Out of bounds")
                        continue
                    
                    url = row[col_idx]
                    content = crawl_url(url, div_class)
                    # Use header name if available, else "Unnamed Column"
                    col_name = header[col_idx] if col_idx < len(header) else f"Column {col_idx}"
                    
                    print(f"  {col_name} (class: {div_class}):")
                    print(f"    URL: {url}")
                    # Truncate long content for readability
                    print(f"    Content: {content[:100]}..." if len(content) > 100 else f"    Content: {content}")

if __name__ == "__main__":
    main()

Key Improvements Explained

  • Flexible Mapping: The column_class_mapping dictionary lets you easily add/remove column-class pairs without rewriting core logic. Just update this dict to match your needs.
  • Robust Error Handling: Added checks for empty URLs, HTTP errors, and out-of-bounds columns to prevent crashes and give clear feedback.
  • Readable Output: Includes row numbers, column names (from your CSV header), and truncated content to make results easier to scan.
  • Scalable: Works with any number of columns and class pairs—no hardcoded limits.

Example urls.csv Input

ID,Product Image URL,Thumbnail URL,Product Description,Specs URL,Manual URL
1,https://example.com/product1/img,https://example.com/product1/thumb,"Wireless Headphones",https://example.com/product1/specs,https://example.com/product1/manual
2,https://example.com/product2/img,https://example.com/product2/thumb,"Bluetooth Speaker",https://example.com/product2/specs,https://example.com/product2/manual

Quick Setup Notes

  1. Install required packages if you haven’t already:
    pip install requests beautifulsoup4
    
  2. Adjust the timeout value in requests.get() if you need longer wait times for slow URLs.
  3. If you want to save results to a new CSV instead of printing, add a csv.writer object and write rows with the collected content.

内容的提问来源于stack exchange,提问作者AnotherUser31

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 09:20:28