You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python新手求助:整合两个API调用并合并DataFrame

Hey there! Let's walk through how to fix this and get your two code snippets working together seamlessly as functions. Since each part works on its own, the main issue is probably how we're passing data between them and structuring the functions to play nice. Here's a step-by-step solution:

Step 1: Turn Document Fetching into a Reusable Function

First, we'll refactor your first code into a function that fetches the 100 documents, converts the JSON response to a pandas DataFrame, and returns it. We'll make sure it keeps the authorID column since that's our link to the author data.

import pandas as pd
import requests
import time  # For rate limiting later

def fetch_documents():
    # Replace this with your actual document API endpoint
    docs_api_url = "https://your-document-api.com/docs?limit=100"
    
    try:
        response = requests.get(docs_api_url)
        response.raise_for_status()  # Raise an error if the API request fails
        docs_json = response.json()
        
        # Adjust the key below to match your API's JSON structure (e.g., "items" instead of "documents")
        docs_df = pd.DataFrame(docs_json["documents"])
        return docs_df
    except requests.exceptions.RequestException as e:
        print(f"Failed to fetch documents: {str(e)}")
        return pd.DataFrame()  # Return empty DataFrame on failure

Step 2: Create a Function to Fetch Author Details

Next, we'll wrap your second code into a function that takes an authorID as input, hits the author API, and returns the author's details. We'll add basic error handling to avoid breaking the whole process if one author's data fails to load.

def fetch_single_author(author_id):
    # Replace this with your actual author API endpoint, inserting the authorID
    author_api_url = f"https://your-author-api.com/authors/{author_id}"
    
    try:
        response = requests.get(author_api_url)
        response.raise_for_status()
        return response.json()
    except requests.exceptions.RequestException as e:
        print(f"Warning: Could not fetch data for author ID {author_id}: {str(e)}")
        return None  # Return None if the request fails

Step 3: Integrate Functions and Merge DataFrames

Now we'll tie everything together: fetch the documents, extract all authorIDs, fetch their corresponding author data, then merge the two DataFrames into one. We'll also add a small delay between author API calls to avoid hitting rate limits (super common with APIs!).

def merge_docs_and_authors():
    # Step 1: Get the documents DataFrame
    documents_df = fetch_documents()
    if documents_df.empty or "authorID" not in documents_df.columns:
        print("Error: No documents found or missing 'authorID' column.")
        return
    
    # Step 2: Fetch author details for every unique authorID
    # Add a 1-second delay between requests to avoid rate limits
    author_details = []
    for auth_id in documents_df["authorID"].unique():  # Fetch unique authors once to save requests
        time.sleep(1)
        author_data = fetch_single_author(auth_id)
        if author_data:
            author_details.append(author_data)
    
    # Convert author data to a DataFrame
    authors_df = pd.DataFrame(author_details)
    # Make sure the author ID column matches the one in documents_df (rename if needed)
    if "id" in authors_df.columns and "authorID" not in authors_df.columns:
        authors_df = authors_df.rename(columns={"id": "authorID"})
    
    # Step 3: Merge the two DataFrames
    # Use "left" merge to keep all documents, even if author data is missing
    merged_df = pd.merge(documents_df, authors_df, on="authorID", how="left")
    
    # Optional: Save the merged data to a CSV for easy access
    merged_df.to_csv("merged_docs_authors.csv", index=False)
    print("Success! Merged data saved to merged_docs_authors.csv")
    return merged_df

# Run the full process when the script is executed
if __name__ == "__main__":
    final_merged_data = merge_docs_and_authors()

Quick Tips to Avoid Headaches

  • Rate Limits: The time.sleep(1) adds a 1-second pause between author API calls. Adjust this if your API allows faster requests, or add retry logic for failed calls.
  • JSON Structure: Double-check the key names in the API responses (like "documents" in the first function) to match what your actual API returns—this is a super common gotcha!
  • Error Handling: We added basic try/except blocks, but you could expand this to log errors to a file instead of just printing them, if needed.
  • Merge Type: We used how="left" to keep all documents even if we couldn't fetch their author details. Use how="inner" if you only want documents with valid author data.

内容的提问来源于stack exchange,提问作者Mamtha Pillai

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:28:15