You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从BeautifulSoup的ResultSet中提取NFL Combine训练数据字段?

Extracting NFL Combine Stats from BeautifulSoup Results

Got it, let's walk through how to pull the drill names and numerical results from your Combinestats list cleanly.

First, let's recap the structure you're working with: each entry in Combinestats is a two-item list. The first item is a string with the drill name, and the second is a BeautifulSoup span element containing the result (with units).

Step-by-Step Solution

Here's a straightforward way to parse each entry and store the data in a usable dictionary:

import urllib.request
from bs4 import BeautifulSoup

# Fix the URL with proper protocol, urllib needs this to work
URL2 = 'https://www.nfl.com/player/deandrewwhite/2552657/combine'
# Specify the HTML parser for consistent behavior across environments
soupCombine = BeautifulSoup(urllib.request.urlopen(URL2), 'html.parser')
Combinestats = soupCombine.find_all("div", attrs={"class": "tp-title"})

# Initialize a dictionary to store cleaned, usable stats
combine_stats = {}

for stat_entry in Combinestats:
    try:
        # Extract and clean the drill name (trim any extra whitespace)
        drill_name = stat_entry[0].strip()
        
        # Pull the result text from the span, then isolate the numerical value
        result_text = stat_entry[1].get_text(strip=True)
        # Split the text to separate number and units, convert the number to a float
        drill_result = float(result_text.split()[0])
        
        # Add the parsed data to our dictionary
        combine_stats[drill_name] = drill_result
    except (IndexError, ValueError) as e:
        # Handle unexpected formatting without crashing the script
        print(f"Skipping malformed stat entry: {stat_entry} | Error: {e}")

# Example of accessing the parsed data
print(combine_stats['3 Cone Drill'])  # Output: 6.97
print(combine_stats['40 Yard Dash'])  # Output: 4.44

Key Details Explained

  • URL Fix: I added https:// to your URL2—urllib will throw an error if you don't include the protocol.
  • Parser Specification: Adding 'html.parser' to BeautifulSoup ensures consistent behavior regardless of your system setup.
  • Text Cleaning: Using .strip() removes accidental leading/trailing whitespace from the drill name.
  • Numeric Extraction: Splitting the result text (e.g., turning "6.97 secs" into ["6.97", "secs"]) lets us grab just the numerical value and convert it to a float for easy integration with other data.
  • Error Handling: The try-except block prevents your script from crashing if a stat entry is missing or formatted unexpectedly.

Once you run this, combine_stats will be a dictionary mapping each drill name to its numerical result—perfect for merging with your existing player list data.

内容的提问来源于stack exchange,提问作者DeeeeRoy

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 03:56:07