如何在网页爬取中生成指定嵌套列表的目标字典结构
网页爬取数据格式整理问题
问题说明
我在网页爬取过程中,需要将获取到的荣誉数据整理成「荣誉标题对应日期+俱乐部列表」的字典格式,现有代码和输出如下,求修改方案。
现有代码
for i in titles: title = i.css('tr[class="bg_Sturm"] > td[class="hauptlink"]::text').get() if title is None: try: date = i.css('tr > td[class="erfolg_table_saison zentriert"] ::text ').get(default = "") club = i.css('tr > td[class="no-border-links"]>a ::text ').get(default = "").strip() if date or club: print({date:club}) except (KeyError, AttributeError): pass else: print(title)
当前输出
2x Champions League participant {'2021': 'Borussia Dortmund'} {'2020': 'Red Bull Salzburg'} 1x German cup winner {'20/21': 'Borussia Dortmund'} 2x Young player of the year {'2020': ''} {'2018': 'Eliteserien'} 1x German Bundesliga runner-up {'19/20': 'Borussia Dortmund'} 3x Footballer of the Year {'2021': 'Norway'} {'2020': 'Norway'} {'2019': 'Austria'} 2x Striker of the Year {'21/22': 'Borussia Dortmund'} {'20/21': 'Borussia Dortmund'} 1x Austrian cup winner {'18/19': 'Red Bull Salzburg'} 3x Top scorer {'20/21': 'UEFA Nations League B'} {'20/21': 'UEFA Champions League'} {'18/19': 'U-20 World Cup 2019'} 1x TM-Player of the season {'2020': 'Austria'}
期望输出格式
希望生成如下结构的字典列表,每个荣誉标题对应一个包含日期和俱乐部的字典列表:
[ { "2次欧冠参赛经历": [ {"date": "2021", "club": "多特蒙德"}, {"date": "2020", "club": "萨尔茨堡红牛"} ] }, { "1次德国杯冠军": [ {"date": "20/21", "club": "多特蒙德"} ] }, { "2次年度最佳年轻球员": [ {"date": "2020", "club": ""}, {"date": "2018", "club": "挪威超级联赛"} ] } # 以此类推 ]
修改后的代码实现
# 初始化结果列表和当前荣誉相关变量 result = [] current_honor = None current_entries = [] for i in titles: title = i.css('tr[class="bg_Sturm"] > td[class="hauptlink"]::text').get() if title is not None: # 若已有未保存的荣誉数据,先存入结果列表 if current_honor is not None and current_entries: result.append({current_honor: current_entries}) # 更新当前荣誉标题,重置条目列表 current_honor = title current_entries = [] else: try: date = i.css('tr > td[class="erfolg_table_saison zentriert"] ::text ').get(default = "") club = i.css('tr > td[class="no-border-links"]>a ::text ').get(default = "").strip() if date or club: # 将日期和俱乐部整理成指定格式加入当前条目 current_entries.append({"date": date, "club": club}) except (KeyError, AttributeError): pass # 循环结束后,添加最后一组荣誉数据 if current_honor is not None and current_entries: result.append({current_honor: current_entries}) # 输出最终结果 print(result)
代码说明
- 用
result存储最终的字典列表,current_honor记录当前处理的荣誉标题,current_entries存储该荣誉对应的所有日期+俱乐部条目。 - 遇到新的荣誉标题时,先将之前的荣誉数据存入结果,再更新当前荣誉并重置条目列表。
- 抓取到日期和俱乐部数据时,整理成
{"date": ..., "club": ...}的格式加入当前条目列表。 - 循环结束后,处理最后一组未存入结果的荣誉数据。
内容的提问来源于stack exchange,提问作者Baraa Zaid
相关产品推荐
相关产品推荐

