You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Oddsportal足球赛事数据提取:CSS选择器修正求助

问题

我尝试从足球赛事网站提取数据,编写了如下Python代码:

def generate_matches(pgSoup, defaultVal=None):
    evtSel = {
        'time': 'div.main-row p.flex',
        'game': 'div.main-row a[title]',
        'score': 'a a[title]+div:has(+a[title])',
        'home_odds': 'a:has(a[title])~div:not(.hidden)',
        'draw_odds': 'a:has(a[title])~div:not(.hidden)+div:nth-last-of-type(3)',
        'away_odds': 'a:has(a[title])~div:nth-last-of-type(2)',
    }

    events, current_group = [], {}
    pgDate = pgSoup.select_one('h1.title[id="next-matches-h1"]')
    if pgDate: pgDate = pgDate.get_text().split(',', 1)[-1].strip()
    for evt in pgSoup.select('div[set]>div:last-child'):
        if evt.parent.select(f':scope>div:first-child+div+div'):
            cgVals = [v.get_text(' ').strip() if v else defaultVal for v in [
                evt.parent.select_one(s) for s in
                [':scope>div:first-child+div>div:first-child',
                 ':scope>div:first-child>a:nth-of-type(2):nth-last-of-type(2)',
                 ':scope>div:first-child>a:nth-of-type(3):last-of-type']]]
            current_group = dict(zip(['date', 'country', 'league'], cgVals))
            if pgDate: current_group['date'] = pgDate

        evtRow = {'date': current_group.get('date', defaultVal)}

        for k, v in evtSel.items():
            v = evt.select_one(v).get_text(' ') if evt.select_one(v) else defaultVal
            evtRow[k] = ' '.join(v.split()) if isinstance(v, str) else v
        evtTeams = evt.select('a div>a[title]')
        evtRow['game'] = ' – '.join(a['title'] for a in evtTeams)
        evtRow['country'] = current_group.get('country', defaultVal)
        evtRow['league'] = current_group.get('league', defaultVal)

        events.append(evtRow)
    return events

当前提取出的DataFrame中time、game、score、home_odds、draw_odds、away_odds字段均为NaN:

date  time game  score  home_odds  draw_odds  away_odds                  country                          league
0     25 July 2023   NaN         NaN        NaN        NaN        NaN                Argentina                  Reserve League
1     25 July 2023   NaN         NaN        NaN        NaN        NaN                Argentina                  Reserve League
2     25 July 2023   NaN         NaN        NaN        NaN        NaN                Argentina                  Reserve League
3     25 July 2023   NaN         NaN        NaN        NaN        NaN                Argentina                  Reserve League
4     25 July 2023   NaN         NaN        NaN        NaN        NaN                   Norway            Division 3 - Group 6

期望得到如下格式的完整数据:

Unnamed: 0         date   time                                  game  score home_odds draw_odds away_odds  country   league
0           0  01 Apr 2023  01:00  Widad Adabi de Boufarik – Temouchent  1 – 1      2.38      2.93      3.06  Algeria  Ligue 2
1           1  01 Apr 2023  01:00                   Relizane – Oued Sly  1 – 2     10.02      5.00      1.28  Algeria  Ligue 2

请问需要使用哪些正确的CSS选择器才能填充所有数据行?

解决方案

原选择器失效是因为目标网站的DOM结构已更新,以下是适配当前页面结构的正确CSS选择器及修改后的代码:

修正后的核心CSS选择器

evtSel = {
    'time': 'div.match-time',
    'game': 'div.match-participants a',
    'score': 'div.match-score',
    'home_odds': 'div.odds-nowrp:nth-child(1)',
    'draw_odds': 'div.odds-nowrp:nth-child(2)',
    'away_odds': 'div.odds-nowrp:nth-child(3)',
}

修改后的完整函数

def generate_matches(pgSoup, defaultVal=None):
    evtSel = {
        'time': 'div.match-time',
        'game': 'div.match-participants a',
        'score': 'div.match-score',
        'home_odds': 'div.odds-nowrp:nth-child(1)',
        'draw_odds': 'div.odds-nowrp:nth-child(2)',
        'away_odds': 'div.odds-nowrp:nth-child(3)',
    }

    events, current_group = [], {}
    # 修正全局日期选择器
    pgDate = pgSoup.select_one('h2.title')
    if pgDate:
        pgDate = pgDate.get_text().strip()

    # 修正赛事行选择器,直接匹配每个赛事的容器
    for evt in pgSoup.select('div.event__match'):
        # 从赛事行的前置头部提取分组信息(日期、国家、联赛)
        group_header = evt.find_previous('div', class_='event__header')
        if group_header:
            date_text = group_header.select_one('div.event__title--date').get_text().strip() if group_header.select_one('div.event__title--date') else defaultVal
            country_link = group_header.select_one('a.event__title--country')
            country_text = country_link.get_text().strip() if country_link else defaultVal
            league_link = group_header.select_one('a.event__title--name')
            league_text = league_link.get_text().strip() if league_link else defaultVal
            current_group = {
                'date': date_text,
                'country': country_text,
                'league': league_text
            }
            # 用全局日期覆盖分组日期(如果存在)
            if pgDate:
                current_group['date'] = pgDate

        evtRow = {'date': current_group.get('date', defaultVal)}

        # 提取各字段数据
        for k, v in evtSel.items():
            elem = evt.select_one(v)
            if elem:
                text = elem.get_text(' ').strip()
                evtRow[k] = ' '.join(text.split())
            else:
                evtRow[k] = defaultVal

        # 拼接球队名称
        teams = evt.select(evtSel['game'])
        if teams:
            evtRow['game'] = ' – '.join([team.get_text().strip() for team in teams])
        else:
            evtRow['game'] = defaultVal

        evtRow['country'] = current_group.get('country', defaultVal)
        evtRow['league'] = current_group.get('league', defaultVal)

        events.append(evtRow)
    return events

关键修正说明

  • 替换了所有失效的CSS选择器,适配网站当前使用的类名(如match-time、event__match等)
  • 重构分组信息提取逻辑,通过赛事行的前置头部元素获取日期、国家和联赛,避免复杂的嵌套选择器
  • 简化字段提取逻辑,直接通过明确的类名定位元素,提升稳定性

内容的提问来源于stack exchange,提问作者PyNoob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 09:43:23