You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在网页爬取中生成指定嵌套列表的目标字典结构

网页爬取数据格式整理问题

问题说明

我在网页爬取过程中,需要将获取到的荣誉数据整理成「荣誉标题对应日期+俱乐部列表」的字典格式,现有代码和输出如下,求修改方案。

现有代码

for i in titles:                                                                                       
    title  = i.css('tr[class="bg_Sturm"] > td[class="hauptlink"]::text').get()                        
    if title is None:                                                                                                                                                                                                                                                                       
        try:                                                                                                                                                                                              
            date = i.css('tr > td[class="erfolg_table_saison zentriert"] ::text ').get(default = "")  
            club = i.css('tr > td[class="no-border-links"]>a ::text ').get(default = "").strip()                                                                                                        
            if date or club:                                                               
                print({date:club})                                                                                                                                                                          
        except (KeyError, AttributeError):                                                            
            pass                                                                                      
    else:                                                                                             
        print(title)   

当前输出

2x Champions League participant
{'2021': 'Borussia Dortmund'}
{'2020': 'Red Bull Salzburg'}
1x German cup winner
{'20/21': 'Borussia Dortmund'}
2x Young player of the year
{'2020': ''}
{'2018': 'Eliteserien'}
1x German Bundesliga runner-up
{'19/20': 'Borussia Dortmund'}
3x Footballer of the Year
{'2021': 'Norway'}
{'2020': 'Norway'}
{'2019': 'Austria'}
2x Striker of the Year
{'21/22': 'Borussia Dortmund'}
{'20/21': 'Borussia Dortmund'}
1x Austrian cup winner
{'18/19': 'Red Bull Salzburg'}
3x Top scorer
{'20/21': 'UEFA Nations League B'}
{'20/21': 'UEFA Champions League'}
{'18/19': 'U-20 World Cup 2019'}
1x TM-Player of the season
{'2020': 'Austria'}

期望输出格式

希望生成如下结构的字典列表,每个荣誉标题对应一个包含日期和俱乐部的字典列表:

[
    {
        "2次欧冠参赛经历": [
            {"date": "2021", "club": "多特蒙德"},
            {"date": "2020", "club": "萨尔茨堡红牛"}
        ]
    },
    {
        "1次德国杯冠军": [
            {"date": "20/21", "club": "多特蒙德"}
        ]
    },
    {
        "2次年度最佳年轻球员": [
            {"date": "2020", "club": ""},
            {"date": "2018", "club": "挪威超级联赛"}
        ]
    }
    # 以此类推
]

修改后的代码实现

# 初始化结果列表和当前荣誉相关变量
result = []
current_honor = None
current_entries = []

for i in titles:                                                                                       
    title = i.css('tr[class="bg_Sturm"] > td[class="hauptlink"]::text').get()                        
    if title is not None:
        # 若已有未保存的荣誉数据,先存入结果列表
        if current_honor is not None and current_entries:
            result.append({current_honor: current_entries})
        # 更新当前荣誉标题,重置条目列表
        current_honor = title
        current_entries = []
    else:                                                                                                                                                                                              
        try:
            date = i.css('tr > td[class="erfolg_table_saison zentriert"] ::text ').get(default = "")  
            club = i.css('tr > td[class="no-border-links"]>a ::text ').get(default = "").strip()                                                                                                        
            if date or club:
                # 将日期和俱乐部整理成指定格式加入当前条目
                current_entries.append({"date": date, "club": club})
        except (KeyError, AttributeError):                                                            
            pass
# 循环结束后,添加最后一组荣誉数据
if current_honor is not None and current_entries:
    result.append({current_honor: current_entries})

# 输出最终结果
print(result)

代码说明

  1. 用result存储最终的字典列表,current_honor记录当前处理的荣誉标题,current_entries存储该荣誉对应的所有日期+俱乐部条目。
  2. 遇到新的荣誉标题时,先将之前的荣誉数据存入结果,再更新当前荣誉并重置条目列表。
  3. 抓取到日期和俱乐部数据时,整理成{"date": ..., "club": ...}的格式加入当前条目列表。
  4. 循环结束后,处理最后一组未存入结果的荣誉数据。

内容的提问来源于stack exchange,提问作者Baraa Zaid

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.19 20:26:16