You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

同class名多HTML块场景下如何仅解析第一个块?BeautifulSoup实现方案

实现方案

两种实现方式可以选,根据你的页面实际结构调整即可:

方案1:取匹配元素列表的首个值

如果确认「今日」模块永远是页面中所有standard-box standard-list类元素的第一个,直接取find_all返回列表的下标0即可,无需遍历所有元素。

注意:原代码缺少BeautifulSoup导入语句,修正后的代码已补上。

import requests
from bs4 import BeautifulSoup

url_news = "https://www.123.org/"
response = requests.get(url_news)
soup = BeautifulSoup(response.content, "html.parser")
items = soup.find_all("div", class_="standard-box standard-list")
news_info = []
# 仅处理第一个匹配的模块
if items:
    item = items[0]
    news_info.append({
        "title": item.find("div", class_="newstext").text,
        "link": item.find("a", class_="newsline article").get("href")
    })

方案2:通过「今日」文本定位(更稳妥)

如果后续页面可能调整模块顺序,推荐先定位包含「今日」文本的节点,再查找对应的目标模块,不受元素顺序影响。
假设你的页面结构参考如下:

<h3>今日</h3>
<div class="standard-box standard-list">今日新闻内容</div>
<h3>昨日</h3>
<div class="standard-box standard-list">昨日新闻内容</div>

可以通过find_next_sibling匹配目标模块,代码如下:

import requests
from bs4 import BeautifulSoup

url_news = "https://www.123.org/"
response = requests.get(url_news)
soup = BeautifulSoup(response.content, "html.parser")
# 按页面实际标签修改,如「今日」放在span标签里就把h3改成span
today_title = soup.find("h3", string="今日")
if today_title:
    item = today_title.find_next_sibling("div", class_="standard-box standard-list")
    news_info = []
    news_info.append({
        "title": item.find("div", class_="newstext").text,
        "link": item.find("a", class_="newsline article").get("href")
    })

内容的提问来源于stack exchange,提问作者makim

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.30 22:15:06