Python提取HTML指定数据异常:无勾选标记时输出缺失排查
HTML数据提取优化:处理有无勾选标记的统一输出问题
需求与问题
提取需求
- 日期(Date):从
div.date.details的title属性中提取日期部分 - 时间(Time):从上述
title属性中提取时间部分 - 第一个Jogo时间:从包含
⚽ Jogos:的文本行中提取第一个有效时间值 - 勾选标记位置:若存在✅,返回其对应的Jogo时间位置(1-4);无标记时返回0
原代码问题
原代码仅在存在✅标记时能正常输出全部数据,无标记时会因index('✅')抛出异常,且仅输出日期和时间,无法返回Jogo 1和Checkmarked:0。
原代码
from bs4 import BeautifulSoup html = """ <div class="body"> <div class="pull_right date details" title="09.03.2023 01:08:10 UTC-03:00"> 01:08 </div> <div class="from_name"> 🤖🥇 𝑬𝒂𝒔𝒚 𝑩𝒐𝒕 - 𝑶𝒗𝒆𝒓 2.5 </div> <div class="text"> Easy Bot - Over 2.5<br><br>🏆 Liga: Sul-Americana<br>🚦 Entrada: Over 2.5 FT<br>⚽ Jogos: ✅ 04:10 04:13 04:16 (04:19)<br><br><strong>Link: </strong><a href="https://www.bet365.com/#/AVR/B146/R%5E1/">https://www.bet365.com/#/AVR/B146/R%5E1/</a><br><br>🍀 24h:100% de acerto nas últimas 24h<br><br>✅✅✅✅✅✅ . </div> </div> """ # parse the HTML soup = BeautifulSoup(html, 'html.parser') d_str = soup.select_one('div.date.details')['title'] calendar = d_str.split(" ") print("Date: ",calendar[0]) print("Time: ",calendar[1]) for sts in soup.select_one('div.text').stripped_strings: if "⚽ Jogos: " in sts: jugos = (sts.split('⚽ Jogos: ')[1].split(" ")) ind = jugos.index('✅')+1 jugos.remove("✅") print("Jogo 1: ", jugos[0]) print("Checkmarked: ", ind)
优化后的代码
from bs4 import BeautifulSoup html = """ <div class="body"> <div class="pull_right date details" title="09.03.2023 01:08:10 UTC-03:00"> 01:08 </div> <div class="from_name"> 🤖🥇 𝑬𝒂𝒔𝒚 𝑩𝒐𝒕 - 𝑶𝒗𝒆𝒓 2.5 </div> <div class="text"> Easy Bot - Over 2.5<br><br>🏆 Liga: Sul-Americana<br>🚦 Entrada: Over 2.5 FT<br>⚽ Jogos: 04:10 04:13 04:16 (04:19)<br><br><strong>Link: </strong><a href="https://www.bet365.com/#/AVR/B146/R%5E1/">https://www.bet365.com/#/AVR/B146/R%5E1/</a><br><br>🍀 24h:100% de acerto nas últimas 24h<br><br>✅✅✅✅✅✅ . </div> </div> """ soup = BeautifulSoup(html, 'html.parser') # 提取日期和时间 d_str = soup.select_one('div.date.details')['title'] calendar = d_str.split(" ") print("Date: ", calendar[0]) print("Time: ", calendar[1]) # 处理Jogos部分 jogo1 = "" checkmarked = 0 for sts in soup.select_one('div.text').stripped_strings: if "⚽ Jogos: " in sts: jugos = sts.split('⚽ Jogos: ')[1].split(" ") # 过滤空字符串,避免分割后的空元素干扰 jugos = [item for item in jugos if item] if '✅' in jugos: checkmarked = jugos.index('✅') + 1 jugos.remove('✅') # 提取第一个有效时间(排除带括号的备用时间) for item in jugos: if not item.startswith('('): jogo1 = item break break # 确保输出,无论有无标记 print("Jogo 1: ", jogo1) print("Checkmarked: ", checkmarked)
关键修改说明
- 初始化默认值:提前定义
jogo1和checkmarked的默认值,确保即使没有标记也能输出结果 - 异常避免:先判断
✅是否在列表中,再执行索引和移除操作,避免抛出ValueError - 过滤无效元素:分割后过滤空字符串,同时跳过带括号的备用时间,确保提取的是有效Jogo时间
- 强制输出:循环结束后统一输出Jogo1和Checkmarked的值,保证无论是否找到标记都能返回完整结果
内容的提问来源于stack exchange,提问作者Marco Almeida
相关产品推荐
相关产品推荐

