基于CSV列循环调用YouTube API批量获取视频评论数据的问题
Looks like you ran into two main issues here: your friend's function only returns the last row of your CSV, and you're not sure how to loop through all songs to call the YouTube API automatically. Let's break this down step by step.
First, let's diagnose the problem with that function: the return statement is placed outside the for loop. That means the loop runs through every row, overwriting artist and title each time, and then returns only the last pair once the loop finishes. To fix this, we can either restructure the function to yield each song's info, or (even better) integrate the loop directly with your API call logic.
Here's a complete, working solution that reads your CSV, loops through each song, calls the YouTube API, and handles data processing:
Full Working Code
import pandas as pd from googleapiclient.discovery import build # Replace with your actual YouTube API key YOUTUBE_API_KEY = "your_api_key_here" service = build('youtube', 'v3', developerKey=YOUTUBE_API_KEY) # Load your song metadata song_metadata = pd.read_csv("metadata.csv") # Loop through every song in the CSV for index, row in song_metadata.iterrows(): song_title = row['title'] artist_name = row['artist'] # Build a more precise search query (adds artist + "official" to avoid irrelevant results) search_query = f"{song_title} {artist_name} official music video" print(f"Processing: {artist_name} - {song_title}") try: # Call YouTube Search API (your original code, now inside the loop) search_results = service.search().list( part='snippet', q=search_query, order='relevance', maxResults=20, type='video', relevanceLanguage='en', safeSearch='moderate', ).execute() # Extract video IDs and fetch comments (add your existing comment logic here) for item in search_results['items']: video_id = item['id']['videoId'] video_title = item['snippet']['title'] print(f"Found matching video: {video_title} (ID: {video_id})") # Insert your comment-fetching code here, e.g.: # comment_response = service.commentThreads().list( # part='snippet', # videoId=video_id, # maxResults=100, # textFormat='plainText' # ).execute() # Then write comments to your CSV (use mode='a' to append rows) # comment_df = pd.DataFrame(your_comment_data) # comment_df.to_csv('song_comments.csv', mode='a', header=(index==0), index=False) except Exception as e: print(f"Error processing {artist_name} - {song_title}: {str(e)}") continue print("All songs processed successfully!")
Key Improvements & Explanations
- Cleaner Loop: Using
iterrows()lets you directly access each row's data without manually indexing, making the code easier to read and maintain. - Better Search Queries: Combining the song title, artist, and "official music video" ensures you're far more likely to find the correct target video, not covers, remixes, or unrelated content.
- Error Handling: The
try-exceptblock prevents a single failed API call (like a video with disabled comments) from crashing the entire batch, and prints targeted errors so you can debug specific songs. - Batch-Friendly CSV Writing: When saving comments, use
mode='a'to append new rows instead of overwriting the file. Theheader=(index==0)ensures the CSV header is only written once (for the first song processed).
Fixing Your Friend's Function (If You Prefer That Approach)
If you want to keep a separate function to retrieve song info, turn it into a generator with yield instead of return—this lets you iterate through all songs one by one without exiting early:
def get_song_pairs(data): for _, row in data.iterrows(): formatted_title = f'"{row["title"]}"' artist = row["artist"] yield artist, formatted_title # Use the generator like this: for artist, title in get_song_pairs(song_metadata): # Call YouTube API with artist and title here print(f"Working on: {artist} - {title}")
Generators don't exit after the first item, so you'll get every song entry from your CSV.
内容的提问来源于stack exchange,提问作者bguerra

