Python新手求助:整合两个API调用并合并DataFrame
Hey there! Let's walk through how to fix this and get your two code snippets working together seamlessly as functions. Since each part works on its own, the main issue is probably how we're passing data between them and structuring the functions to play nice. Here's a step-by-step solution:
Step 1: Turn Document Fetching into a Reusable Function
First, we'll refactor your first code into a function that fetches the 100 documents, converts the JSON response to a pandas DataFrame, and returns it. We'll make sure it keeps the authorID column since that's our link to the author data.
import pandas as pd import requests import time # For rate limiting later def fetch_documents(): # Replace this with your actual document API endpoint docs_api_url = "https://your-document-api.com/docs?limit=100" try: response = requests.get(docs_api_url) response.raise_for_status() # Raise an error if the API request fails docs_json = response.json() # Adjust the key below to match your API's JSON structure (e.g., "items" instead of "documents") docs_df = pd.DataFrame(docs_json["documents"]) return docs_df except requests.exceptions.RequestException as e: print(f"Failed to fetch documents: {str(e)}") return pd.DataFrame() # Return empty DataFrame on failure
Step 2: Create a Function to Fetch Author Details
Next, we'll wrap your second code into a function that takes an authorID as input, hits the author API, and returns the author's details. We'll add basic error handling to avoid breaking the whole process if one author's data fails to load.
def fetch_single_author(author_id): # Replace this with your actual author API endpoint, inserting the authorID author_api_url = f"https://your-author-api.com/authors/{author_id}" try: response = requests.get(author_api_url) response.raise_for_status() return response.json() except requests.exceptions.RequestException as e: print(f"Warning: Could not fetch data for author ID {author_id}: {str(e)}") return None # Return None if the request fails
Step 3: Integrate Functions and Merge DataFrames
Now we'll tie everything together: fetch the documents, extract all authorIDs, fetch their corresponding author data, then merge the two DataFrames into one. We'll also add a small delay between author API calls to avoid hitting rate limits (super common with APIs!).
def merge_docs_and_authors(): # Step 1: Get the documents DataFrame documents_df = fetch_documents() if documents_df.empty or "authorID" not in documents_df.columns: print("Error: No documents found or missing 'authorID' column.") return # Step 2: Fetch author details for every unique authorID # Add a 1-second delay between requests to avoid rate limits author_details = [] for auth_id in documents_df["authorID"].unique(): # Fetch unique authors once to save requests time.sleep(1) author_data = fetch_single_author(auth_id) if author_data: author_details.append(author_data) # Convert author data to a DataFrame authors_df = pd.DataFrame(author_details) # Make sure the author ID column matches the one in documents_df (rename if needed) if "id" in authors_df.columns and "authorID" not in authors_df.columns: authors_df = authors_df.rename(columns={"id": "authorID"}) # Step 3: Merge the two DataFrames # Use "left" merge to keep all documents, even if author data is missing merged_df = pd.merge(documents_df, authors_df, on="authorID", how="left") # Optional: Save the merged data to a CSV for easy access merged_df.to_csv("merged_docs_authors.csv", index=False) print("Success! Merged data saved to merged_docs_authors.csv") return merged_df # Run the full process when the script is executed if __name__ == "__main__": final_merged_data = merge_docs_and_authors()
Quick Tips to Avoid Headaches
- Rate Limits: The
time.sleep(1)adds a 1-second pause between author API calls. Adjust this if your API allows faster requests, or add retry logic for failed calls. - JSON Structure: Double-check the key names in the API responses (like
"documents"in the first function) to match what your actual API returns—this is a super common gotcha! - Error Handling: We added basic try/except blocks, but you could expand this to log errors to a file instead of just printing them, if needed.
- Merge Type: We used
how="left"to keep all documents even if we couldn't fetch their author details. Usehow="inner"if you only want documents with valid author data.
内容的提问来源于stack exchange,提问作者Mamtha Pillai

