You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用requests_HTML爬取网页时无法获取caption元素文本的问题

解决方案

针对页面未闭合<br>标签导致无法提取<caption>文本的问题,推荐使用BeautifulSoup配合lxml解析器处理不规范HTML,具体修改如下:

1. 依赖安装

先安装所需库:

pip install beautifulsoup4 lxml

2. 修改代码

替换原requests-html的解析逻辑,改用BeautifulSoup处理:

from bs4 import BeautifulSoup

def grab_ranking():
    tournament_list = grab_tournament_metadata()
    for item in tournament_list:
        url_to_scrape = f'https://www.kayak-polo.info/kphistorique.php?Group={item[1]}&lang=en'
        response = session.get(url_to_scrape)
        print(url_to_scrape)
        
        # 用lxml解析器处理不规范HTML,容错性更强
        soup = BeautifulSoup(response.content, 'lxml')
        season_data = soup.select('body > div.container-fluid > div > article')
        
        for season in season_data:
            # 提取赛季年份
            season_year_raw = season.select_one('h3 > div.col-md-6.col-sm-6').get_text(strip=True)
            season_year = season_year_raw.replace('Season ', '')
            print(season_year)

            # 定位分类容器并提取caption文本
            category_table = season.select_one('div.col-md-3.col-sm-6.col-xs-12')
            if category_table:
                umbrella_competition_name = category_table.select_one('caption').get_text(strip=True)
                competition_name = umbrella_competition_name + " " + season_year
                # 可添加后续存储/处理逻辑
                print(competition_name)

原理说明

原requests-html默认解析器对未闭合标签的容错性较差,会导致DOM结构解析混乱,无法正确定位<caption>元素。而lxml解析器专门针对不规范HTML做了优化,能正确构建DOM树,确保元素定位和文本提取正常工作。

内容的提问来源于stack exchange,提问作者Jesper van Beemdelust

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.14 19:50:26