使用requests_HTML爬取网页时无法获取caption元素文本的问题
解决方案
针对页面未闭合<br>标签导致无法提取<caption>文本的问题,推荐使用BeautifulSoup配合lxml解析器处理不规范HTML,具体修改如下:
1. 依赖安装
先安装所需库:
pip install beautifulsoup4 lxml
2. 修改代码
替换原requests-html的解析逻辑,改用BeautifulSoup处理:
from bs4 import BeautifulSoup def grab_ranking(): tournament_list = grab_tournament_metadata() for item in tournament_list: url_to_scrape = f'https://www.kayak-polo.info/kphistorique.php?Group={item[1]}&lang=en' response = session.get(url_to_scrape) print(url_to_scrape) # 用lxml解析器处理不规范HTML,容错性更强 soup = BeautifulSoup(response.content, 'lxml') season_data = soup.select('body > div.container-fluid > div > article') for season in season_data: # 提取赛季年份 season_year_raw = season.select_one('h3 > div.col-md-6.col-sm-6').get_text(strip=True) season_year = season_year_raw.replace('Season ', '') print(season_year) # 定位分类容器并提取caption文本 category_table = season.select_one('div.col-md-3.col-sm-6.col-xs-12') if category_table: umbrella_competition_name = category_table.select_one('caption').get_text(strip=True) competition_name = umbrella_competition_name + " " + season_year # 可添加后续存储/处理逻辑 print(competition_name)
原理说明
原requests-html默认解析器对未闭合标签的容错性较差,会导致DOM结构解析混乱,无法正确定位<caption>元素。而lxml解析器专门针对不规范HTML做了优化,能正确构建DOM树,确保元素定位和文本提取正常工作。
内容的提问来源于stack exchange,提问作者Jesper van Beemdelust
相关产品推荐
相关产品推荐

