You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python/BeautifulSoup抓取同名标签中的指定机构信息

Fixing Your BeautifulSoup Institution Scraping Issue

Hey there! I see where the problem is with your current code—you're not targeting the specific <dd> that corresponds to the "Institution" <dt> within each course block. Let's get this sorted out.

Why Your Current Code Isn't Working

Your line:

institutions = [x.text.strip() for x in bsObj.find_all('div', 'course-meta col-sm-12', 'dd')]

is actually grabbing all text inside every div.course-meta col-sm-12 element, not just the <dd> tied to "Institution". The third argument in find_all() isn't for specifying a child tag like that—it's used for other parameters like recursive.

The Solution: Target Specific <dt> + <dd> Pairs

We need to loop through each course's metadata block, find the <dt> that says "Institution", then grab its immediately following <dd> sibling. Here's how to do it:

Option 1: Explicit Loop (Easy to Read)

from urllib.request import urlopen
from bs4 import BeautifulSoup

html = urlopen("http://www.oeconsortium.org/courses/search/?search=statistics")
bsObj = BeautifulSoup(html.read(), "html.parser")

institutions = []

# Iterate through each course's metadata container
for course_meta in bsObj.find_all('div', class_='course-meta col-sm-12'):
    # Look for all <dt> tags in this container
    for dt in course_meta.find_all('dt'):
        # Check if this <dt> is the "Institution" label
        if dt.text.strip() == "Institution":
            # Grab the next sibling <dd> (the actual institution value)
            institution_dd = dt.find_next_sibling('dd')
            if institution_dd:  # Make sure the <dd> exists to avoid errors
                institutions.append(institution_dd.text.strip())
            break  # No need to check other <dt>s once we find the right one

# Print your results
for inst in institutions:
    print(inst)

Option 2: More Concise Version

If you prefer a tighter approach, you can use a lambda to target the "Institution" <dt> directly:

from urllib.request import urlopen
from bs4 import BeautifulSoup

html = urlopen("http://www.oeconsortium.org/courses/search/?search=statistics")
bsObj = BeautifulSoup(html.read(), "html.parser")

institutions = []
for course_meta in bsObj.find_all('div', class_='course-meta col-sm-12'):
    # Find the <dt> with text containing "Institution"
    institution_label = course_meta.find('dt', string=lambda t: t and "Institution" in t.strip())
    if institution_label:
        # Get the corresponding <dd>
        institution_value = institution_label.find_next_sibling('dd')
        if institution_value:
            institutions.append(institution_value.text.strip())

print(institutions)

Key Takeaways

  • Always narrow down your scope first: start with each course's container (div.course-meta col-sm-12) before digging into child elements.
  • Use find_next_sibling() to pair <dt> labels with their corresponding <dd> values—this works perfectly because they're adjacent in the HTML structure.
  • Add checks (like if institution_dd:) to handle cases where a course might not have an Institution listed, preventing AttributeErrors.

内容的提问来源于stack exchange,提问作者RichardS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.28 03:48:21