求助:使用BeautifulSoup解析新闻网站时无法识别指定main标签子类
解决BeautifulSoup无法定位指定标签子类的问题
问题背景
需要解析某新闻网站,新闻标题与链接预期位于<main class='single-module__Main-sc-1qdjg1k-0 iMZnZU'>标签内,但使用BeautifulSoup查询该标签子类时失败。soup.prettify()输出的关键HTML结构如下:
<div class="single-module__Inner-sc-1qdjg1k-1 mIFFS"> <header class="sticky-header"> <div data-fusion-collection="features" data-fusion-message="Could not render component [features:global/main-navigation]" data-fusion-type="global/main-navigation" id="f0fj2XGPHYgA9Rb" style="display:none"> </div> </header> <main class="single-module__Main-sc-1qdjg1k-0 iMZnZU"> <div data-fusion-collection="features" data-fusion-message="Could not render component [features:global/search-page]" data-fusion-type="global/search-page" id="f0f1nZqKTkTE1lq" style="display:none"> </div> </main> <footer> <div data-fusion-collection="features" data-fusion-message="Could not render component [features:global/footer]" data-fusion-type="global/footer" id="f0fUuwTNo76j9AD" style="display:none"> </div> </footer>
尝试的查询代码:
q = soup.find('div', class_='layout-container').find('div','single-module__Inner-sc-1qdjg1k-1mIFFS') print(q.find('main','single-module__Main-sc-1qdjg1k-0 iMZnZU')
问题分析
- 动态渲染导致内容缺失:从HTML源码可见,
main标签内仅包含一个display:none的空div,且带有Could not render component提示,说明目标新闻内容是通过JavaScript动态加载的。BeautifulSoup只能解析初始静态HTML,无法获取JS执行后渲染的内容。 - 代码语法错误:
- 第二个
find的class名错误:'single-module__Inner-sc-1qdjg1k-1mIFFS'应为'single-module__Inner-sc-1qdjg1k-1 mIFFS'(两个类名间有空格) - 最后一行
print语句缺少闭合括号)
- 第二个
解决方案
方案1:使用动态渲染工具获取完整页面
使用selenium或playwright模拟浏览器加载页面,等待JS执行完成后再解析源码:
from selenium import webdriver from bs4 import BeautifulSoup from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.common.by import By # 初始化Chrome浏览器(需提前安装对应版本的ChromeDriver) driver = webdriver.Chrome() driver.get("目标新闻网站的URL") # 显式等待目标main标签加载完成(更可靠) wait = WebDriverWait(driver, 15) wait.until(EC.presence_of_element_located((By.CLASS_NAME, "iMZnZU"))) # 获取渲染后的完整页面源码 page_source = driver.page_source driver.quit() # 解析源码 soup = BeautifulSoup(page_source, 'html.parser') main_tag = soup.find('main', class_='single-module__Main-sc-1qdjg1k-0 iMZnZU') # 现在可以正常查询main标签内的子类元素 print(main_tag.prettify())
方案2:直接调用API接口
打开浏览器开发者工具的「Network」面板,筛选「XHR/Fetch」类型请求,找到加载新闻内容的API接口,直接请求该接口获取JSON格式数据,这种方式比爬页面更高效且稳定。
代码语法修正(仅适用于静态页面场景)
若页面为静态内容,修正后的查询代码如下:
q = soup.find('div', class_='layout-container').find('div', class_='single-module__Inner-sc-1qdjg1k-1 mIFFS') print(q.find('main', class_='single-module__Main-sc-1qdjg1k-0 iMZnZU'))
内容的提问来源于stack exchange,提问作者ИНДУС Геймдев
相关产品推荐
相关产品推荐

