You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何将BeautifulSoup抓取的博彩网站数据转换为规整的DataFrame?

解决博彩网站数据抓取并整理为DataFrame的问题

问题分析

你当前代码用BeautifulSoup.find_all获取的是HTML标签对象列表,直接输出会显示原始标签结构导致文本混乱;同时分开提取不同class的元素,容易出现字段对应错位的问题。以下是无需正则表达式的解决方案:

优化步骤

  1. 智能等待页面加载:替换固定延迟time.sleep为条件等待,确保目标元素加载完成
  2. 通过父容器统一提取数据:定位每个比赛的外层容器,在容器内精准提取对应字段,避免数据错位
  3. 清洗拆分文本:对提取的文本做格式化处理,拆分主客场信息

完整代码

import numpy as np
import pandas as pd
from selenium import webdriver
from selenium.webdriver.chrome.options import Options
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By
from bs4 import BeautifulSoup

DRIVER_PATH = 'C:\\executables\\chromedriver.exe'

options = Options()
options.headless = True
options.add_argument("--window-size=1920,1200")

driver = webdriver.Chrome(options=options, executable_path=DRIVER_PATH)
driver.get("https://www.nike.sk/live-stavky/futbal")

# 等待对阵元素加载完成,最长等待20秒
WebDriverWait(driver, 20).until(
    EC.presence_of_element_located((By.CLASS_NAME, 'match-opponents'))
)

soup = BeautifulSoup(driver.page_source, 'html.parser')
driver.quit()  # 用完及时关闭浏览器

# 获取所有比赛的父容器(从对阵元素向上追溯,根据页面DOM结构调整层级)
match_containers = [elem.parent.parent for elem in soup.find_all(class_='ellipsis f-condensed c-black-100 text-extra-bold match-opponents pr-10')]

# 遍历每个容器提取单场比赛数据
match_data = []
for container in match_containers:
    # 提取比赛时间
    time_elem = container.find(class_='ellipsis flex fs-10 c-black-50 justify-between pr-5')
    match_time = time_elem.get_text(strip=True) if time_elem else ""
    
    # 提取并拆分主客场
    opp_elem = container.find(class_='ellipsis f-condensed c-black-100 text-extra-bold match-opponents pr-10')
    opp_text = opp_elem.get_text(strip=True) if opp_elem else ""
    home_team, away_team = opp_text.split(' - ') if ' - ' in opp_text else ("", "")
    
    # 提取比分(兼容两种状态的class)
    score_elem = container.find(class_='flex justify-center text-right flex-col match-score-col fs-12 c-orange text-extra-bold') or \
                 container.find(class_='flex justify-center text-right flex-col match-score-col fs-12 text-extra-bold c-default-light')
    match_score = score_elem.get_text(strip=True) if score_elem else ""
    
    match_data.append({
        "比赛时间": match_time,
        "主队": home_team,
        "客队": away_team,
        "比分": match_score
    })

# 转换为DataFrame
df = pd.DataFrame(match_data)
print(df)

关键说明

  • 智能等待:用WebDriverWait替代固定睡眠,避免因页面加载缓慢导致的空数据问题
  • 父容器提取:通过对阵元素的父容器定位单场比赛,确保时间、对阵、比分一一对应
  • 文本处理:用split(' - ')拆分主客场,无需正则;用get_text(strip=True)清除文本中的多余空格和换行
  • 多状态兼容:通过or逻辑匹配两种状态的比分元素,覆盖进行中、已结束等不同比赛状态

内容的提问来源于stack exchange,提问作者314mip

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.05 11:25:56