You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup解析网页为何仅返回单条title字段的解析结果

问题原因

  • 培养方向提取仅返回单条的核心原因:你使用了唯一ID属性fak_id_7a3586aa7b32182f036c0dab143d2df8_493作为筛选条件,HTML规范中id属性全局唯一,因此find_all最多只能匹配到1个元素
  • 输出结果存在大量空白:提取文本时未做去空格处理
  • 若教育项目也只返回单条,说明当前的外层容器筛选逻辑find_all('div', class_= 'column-center_rasp')只能匹配到单个容器,需要调整选择器直接匹配所有教育项目标题

修复方案

  1. 去掉培养方向的id筛选条件,直接匹配所有方向标题的通用classgrpPeriod
  2. 所有提取的文本调用strip()方法清除前后多余空白
  3. 教育项目提取逻辑调整为直接匹配所有标题节点,避免外层容器限制

修复后代码

import requests
from bs4 import BeautifulSoup

# 这里替换成你自己的HEADERS和URL
HEADERS = {'User-Agent': '替换为你的浏览器UA信息'}
URL = '替换为目标页面地址'

def get_HTML(url, params=None):
    request = requests.get(url, headers=HEADERS, params=params)
    return request

def get_Content(html):
    soup = BeautifulSoup(html, 'html.parser')

    # 教育项目提取逻辑优化
    eduProgram_titles = soup.find_all('div', class_='headerEduPrograms')
    eduProgram = []
    for title in eduProgram_titles:
        eduProgram.append({
            'title': title.get_text().strip()
        })
    print(eduProgram)
    
    # 培养方向提取逻辑优化,去掉唯一id筛选
    eduDirection_titles = soup.find_all('div', class_='grpPeriod')
    eduDirections = []
    for title in eduDirection_titles:
        eduDirections.append({
            'title': title.get_text().strip()
        })
    print(eduDirections)
    

def parse():
    html = get_HTML(URL)
    if html.status_code == 200:
        get_Content(html.text)
    else:
        print('Error')

parse()

补充说明

如果修改后仍存在教育项目提取不全的情况,可右键页面对应学士/硕士标题元素,检查是否属于iframe加载、前端动态渲染的内容,若为动态渲染需要使用Selenium模拟浏览器或者抓包接口的方式提取。

内容的提问来源于stack exchange,提问作者Mono

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.29 21:54:04