使用Python字典从NHL API抓取数据:嵌套数据索引遇阻求助
Hey there! Let's get your fun data scraping project back on track. The core issue here is that you're trying to parse structured JSON data (which translates to nested dictionaries/lists in Python) as CSV, which breaks the nested structure and makes it impossible to index fields like eventTypeId or coordinates. Let's fix this step by step.
What's Wrong With Your Current Code?
The NHL API returns a JSON object, not plain CSV text. When you use game.text and split it into lines for csv.reader, you're turning a organized nested dataset into unstructured, meaningless lines—this is why your indexing attempts aren't working.
Correct Approach to Parse & Filter the Data
Here's how to properly fetch, parse, and extract the shot/goal coordinates you need:
Step 1: Parse the JSON Directly
Use the .json() method of the requests response to convert the API result straight into a Python dictionary. This preserves all the nested structure.
Step 2: Locate the Event Dataset
All game events live in game_data['liveData']['plays']['allPlays']—this is a list of dictionaries, each representing one event (shot, goal, penalty, faceoff, etc.).
Step 3: Filter for Shots & Goals
Loop through each event, check if eventTypeId is either "SHOT" or "GOAL" (note: these are uppercase in the API response), then extract the x and y coordinates. We'll add a safety check in case any event is missing coordinates.
Step 4: Write Results to a File
We'll use the csv module to save the data in a clean, easy-to-analyze format (you can switch to a text file if you prefer, but CSV is better for later data work).
Full Working Code
import requests as rq import csv GAME_ID = "2017021121" # Game ID breakdown: first 4 digits = year, middle 2 = season stage (01=preseason, 02=regular, 03=playoffs, 04=all-star) url = f"https://statsapi.web.nhl.com/api/v1/game/{GAME_ID}/feed/live" def main(): # Fetch and parse the JSON data directly response = rq.get(url) game_data = response.json() # Access all game events from the nested structure all_events = game_data['liveData']['plays']['allPlays'] # Open a CSV file to store coordinates (using 'w' to overwrite, 'a' to append if needed) with open(f'coordinates.{GAME_ID}.csv', 'w', newline='') as csvfile: # Define columns for clarity fieldnames = ['event_type', 'x_coordinate', 'y_coordinate', 'period', 'game_time'] writer = csv.DictWriter(csvfile, fieldnames=fieldnames) writer.writeheader() # Loop through each event to filter shots and goals for event in all_events: event_type = event['result']['eventTypeId'] # Only keep shots and goals if event_type in ['SHOT', 'GOAL']: # Safely extract coordinates (handle cases where coords might be missing) coords = event.get('coordinates', {}) x_coord = coords.get('x', 'N/A') y_coord = coords.get('y', 'N/A') # Optional: add extra context like period and game time period = event['about']['period'] game_time = event['about']['periodTime'] # Write the data row to CSV writer.writerow({ 'event_type': event_type, 'x_coordinate': x_coord, 'y_coordinate': y_coord, 'period': period, 'game_time': game_time }) if __name__ == "__main__": main()
Key Notes:
- Case Sensitivity: The API uses uppercase values for
eventTypeId, so we check for['SHOT', 'GOAL']instead of lowercase. - Safety Checks: Using
.get()ensures your code won't crash if an event unexpectedly missing coordinates (though shots and goals almost always have them). - File Flexibility: If you want a plain text file instead of CSV, replace the CSV writer with simple
file.write()calls to format the data as you like.
内容的提问来源于stack exchange,提问作者samdisorbo

