You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python批量处理CSV行数据调用API并更新匹配列的技术实现问询

Solution to Loop Through CSV Rows and Call API

Got it, let's break this down step by step to get your loop working smoothly. Here's how you can scale your single-row code to handle hundreds of entries in a CSV:

1. Import Required Libraries

First, make sure you have the necessary modules installed (you probably already have requests, but csv is part of Python's standard library):

import csv
import requests
import json
import time  # For rate limiting, if needed

2. Set Up Your API Credentials

Keep your auth token and base URL handy—replace placeholders with your actual values:

TOKEN = "your-auth-token-here"
BASE_URL = "https://your-base-url.com"
HEADERS = {
    'Authorization': TOKEN,
    'Content-Type': 'application/json; charset=utf-8',
}

3. Process the CSV File

We'll use csv.DictReader to read rows as dictionaries (easy to access columns by name) and csv.DictWriter to write the updated rows with the new match column.

Here's the full loop logic:

# Input and output file paths (use different names to avoid overwriting original)
input_csv = "your-input-file.csv"
output_csv = "your-output-file.csv"

# Open input and output files
with open(input_csv, mode='r', newline='', encoding='utf-8') as infile, \
     open(output_csv, mode='w', newline='', encoding='utf-8') as outfile:
    
    # Create reader and writer objects
    reader = csv.DictReader(infile)
    # Add "match" to the list of fieldnames for the output
    fieldnames = reader.fieldnames + ["match"]
    writer = csv.DictWriter(outfile, fieldnames=fieldnames)
    
    # Write the header row to the output CSV
    writer.writeheader()
    
    # Loop through each row in the input CSV
    for row in reader:
        try:
            # Extract data from the row (adjust column names to match your CSV!)
            company_data = {
                "name": row["company_name"],  # Replace with your actual column name
                "email_domain": row["email_domain"],  # Match your CSV column
                "url": row["website_url"]  # Match your CSV column
            }
            
            # Convert the data to JSON string (no more manual quote escaping!)
            request_data = json.dumps([company_data])
            
            # Send POST request to the API
            response = requests.post(
                f"{BASE_URL}/api/match",
                headers=HEADERS,
                data=request_data
            )
            
            # Check if the request was successful and if data was returned
            response.raise_for_status()  # Raise error for HTTP status codes like 400/500
            api_response = response.json()
            
            # Determine if we got a match (adjust this logic to fit your API's response!)
            # Example: If the API returns a non-empty list/dictionary, mark as "yes"
            if api_response:
                row["match"] = "yes"
            else:
                row["match"] = "no"
            
            # Add a small delay to avoid hitting API rate limits (adjust as needed)
            time.sleep(1)
            
        except requests.exceptions.RequestException as e:
            # Handle API errors (timeout, connection issues, HTTP errors)
            print(f"Error processing row {row}: {str(e)}")
            row["match"] = "error"  # Mark rows that failed to process
        except KeyError as e:
            # Handle missing columns in the CSV
            print(f"Missing column {str(e)} in row {row}")
            row["match"] = "missing_data"
        
        # Write the updated row to the output CSV
        writer.writerow(row)

print("Processing complete! Check the output CSV for results.")

Key Notes to Adjust for Your Use Case

  • CSV Column Names: Make sure the keys in company_data (like row["company_name"]) match exactly with the column headers in your input CSV.
  • API Response Logic: The if api_response: check is a simple example. You'll need to adjust this to match how your API indicates a successful match. For example, maybe the API returns a matched boolean field—so you'd use if api_response.get("matched"): instead.
  • Rate Limiting: If your API has rate limits, increase the time.sleep() value (e.g., time.sleep(2) for 2 seconds between requests) to avoid being blocked.
  • Error Handling: The try-except blocks will catch common issues, but you can expand them based on your specific needs (e.g., handling invalid URLs in the CSV).

Testing Tips

  • Test with a small subset of your CSV first (like 2-3 rows) to make sure everything works before processing hundreds of entries.
  • Print the api_response value for a few rows to understand exactly what the API returns, so you can tweak the match logic correctly.

内容的提问来源于stack exchange,提问作者nwc

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 10:17:28