无法用BeautifulSoup定位指定class的div,无法提取Blockworks文章内容
解决Blockworks文章内容提取失败的问题
问题分析
你当前代码用完整的Tailwind CSS长类列表定位div,这类框架生成的样式类极易因页面样式更新、响应式适配而变动,导致无法匹配目标元素。
可行解决方案
1. 简化Class选择器,保留核心标识类
放弃冗长的class列表,只使用最稳定的核心类组合定位:
from bs4 import BeautifulSoup import requests response = requests.get('https://blockworks.co/news/btc-backed-sustainable-token') soup = BeautifulSoup(response.content, 'html.parser') # 仅用核心class组合定位目标容器 div_element = soup.find('div', class_='prose prose-purple') if div_element: div_text = div_element.get_text(strip=True, separator='\n') print(div_text) else: print("未找到目标内容容器")
若担心单类不够精准,可通过lambda匹配包含关键类的元素:
div_element = soup.find('div', class_=lambda c: c and 'prose' in c and 'prose-purple' in c)
2. 结合父元素层级定位
先定位文章主容器(如<article>标签),再在其中查找内容div,层级定位稳定性更强:
from bs4 import BeautifulSoup import requests response = requests.get('https://blockworks.co/news/btc-backed-sustainable-token') soup = BeautifulSoup(response.content, 'html.parser') # 先找文章主容器,再遍历内部内容div article = soup.find('article') if article: div_element = article.find('div', class_='prose') if div_element: print(div_element.get_text(strip=True, separator='\n')) else: print("文章容器内未找到内容div") else: print("未找到文章主容器")
3. 处理动态加载内容
若页面内容由JavaScript动态渲染,requests仅能获取静态HTML,需用Selenium获取渲染后的完整页面:
from bs4 import BeautifulSoup from selenium import webdriver from selenium.webdriver.chrome.options import Options # 配置无头浏览器 options = Options() options.add_argument('--headless=new') driver = webdriver.Chrome(options=options) driver.get('https://blockworks.co/news/btc-backed-sustainable-token') soup = BeautifulSoup(driver.page_source, 'html.parser') div_element = soup.find('div', class_='prose prose-purple') if div_element: print(div_element.get_text(strip=True, separator='\n')) else: print("未找到目标内容容器") driver.quit()
内容的提问来源于stack exchange,提问作者Siqueler
相关产品推荐
相关产品推荐

