求助解析非标准th&tr格式的网页足球预测Div表格
非标准Div结构足球预测表格解析方案
针对目标页面用Div模拟表格的结构,直接用find_all("tr")肯定无效,以下是基于BeautifulSoup的可行解析步骤和代码:
1. 定位表格核心容器
首先从页面源码里找到包裹整个预测表格的外层Div,通过类名定位(目标页面的表格容器类为PredictionTable):
def parse_table(soup): # 定位表格容器 table_container = soup.find("div", class_="PredictionTable") if not table_container: return []
2. 提取表头字段
表头区域的每个字段对应一个带PredictionTable-headerCell类的Div,遍历提取文本作为表头:
# 提取表头 header_cells = table_container.find_all("div", class_="PredictionTable-headerCell") headers = [cell.get_text(strip=True) for cell in header_cells]
3. 遍历提取每行数据
每行数据被包裹在PredictionTable-row类的Div中,每行内的单元格对应PredictionTable-cell类的Div,将单元格文本和表头一一对应:
# 提取所有行数据 rows = table_container.find_all("div", class_="PredictionTable-row") table_data = [] for row in rows: cells = row.find_all("div", class_="PredictionTable-cell") # 确保单元格数量和表头一致,跳过异常行 if len(cells) != len(headers): continue row_data = dict(zip(headers, [cell.get_text(strip=True) for cell in cells])) table_data.append(row_data) return table_data
补充说明
- 如果遇到部分单元格包含嵌套元素(比如球队logo、概率图标),可以通过
cell.find("span")或其他子元素定位文本,避免提取到多余内容 - 若页面有分页,需要结合Selenium的滚动或点击分页按钮,加载完所有数据后再调用解析函数
内容的提问来源于stack exchange,提问作者Michael WS
相关产品推荐
相关产品推荐

