使用BeautifulSoup无法找到指定class的div标签问题求助
解决Zoomit新闻URL爬取为空的问题
你遇到的核心问题是:requests获取的是服务器返回的原始HTML,而目标页面的新闻内容是通过JavaScript动态渲染生成的——那些带动态后缀的class标签(比如flex__Flex-le1v16-0 eQTmR),是浏览器加载JS后才生成的,原始HTML里根本不存在,所以BeautifulSoup自然找不到。另外你用的Instagram移动端UA,也可能导致网站返回适配特殊端的异常内容。
以下是两种可行的解决方案:
方案一:用Selenium模拟浏览器渲染
Selenium可以模拟真实浏览器完整加载页面,等待JS渲染完成后再抓取内容,完美适配动态页面:
from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.support.ui import WebDriverWait from selenium.webdriver.support import expected_conditions as EC from bs4 import BeautifulSoup news_url = [] driver = webdriver.Chrome() # 需要提前安装ChromeDriver并配置环境 try: target_url = 'https://www.zoomit.ir/archive/?sort=Newest&skip=20&publishPeriod=Last24Hours' driver.get(target_url) # 等待新闻容器加载完成,用模糊匹配避免动态class变化 WebDriverWait(driver, 10).until( EC.presence_of_element_located((By.CSS_SELECTOR, "div[class*='flex__Flex']")) ) # 获取渲染后的完整页面源码 page_source = driver.page_source soup = BeautifulSoup(page_source, 'html.parser') # 模糊匹配新闻容器和链接标签 news_containers = soup.find_all('div', class_=lambda x: x and 'flex__Flex' in x) for container in news_containers: link_tag = container.find('a', class_=lambda x: x and 'link__CustomNextLink' in x) if link_tag and 'href' in link_tag.attrs: news_url.append(link_tag['href']) finally: driver.quit() print(news_url)
方案二:直接调用网站API接口(更高效)
打开浏览器开发者工具的Network标签,刷新页面就能发现,Zoomit是通过API接口获取新闻数据的。直接请求API可以拿到结构化JSON,比解析HTML更稳定高效:
import requests news_url = [] headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36' } api_url = 'https://www.zoomit.ir/api/v1/archive' params = { 'sort': 'Newest', 'skip': 20, 'publishPeriod': 'Last24Hours' } response = requests.get(api_url, headers=headers, params=params) if response.status_code == 200: data = response.json() for item in data.get('data', []): news_url.append(f"https://www.zoomit.ir{item['url']}") print(news_url)
额外提示
- 动态生成的class(带随机后缀的)别硬编码,用模糊匹配或者找更稳定的元素特征(比如标签层级、固定属性)。
- 尽量用常规浏览器的User-Agent,避免用特殊APP的UA,减少被网站拦截或返回异常内容的概率。
内容的提问来源于stack exchange,提问作者Aghmehdiata
相关产品推荐
相关产品推荐

