如何用Python/BeautifulSoup抓取同名标签中的指定机构信息
Hey there! I see where the problem is with your current code—you're not targeting the specific <dd> that corresponds to the "Institution" <dt> within each course block. Let's get this sorted out.
Why Your Current Code Isn't Working
Your line:
institutions = [x.text.strip() for x in bsObj.find_all('div', 'course-meta col-sm-12', 'dd')]
is actually grabbing all text inside every div.course-meta col-sm-12 element, not just the <dd> tied to "Institution". The third argument in find_all() isn't for specifying a child tag like that—it's used for other parameters like recursive.
The Solution: Target Specific <dt> + <dd> Pairs
We need to loop through each course's metadata block, find the <dt> that says "Institution", then grab its immediately following <dd> sibling. Here's how to do it:
Option 1: Explicit Loop (Easy to Read)
from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen("http://www.oeconsortium.org/courses/search/?search=statistics") bsObj = BeautifulSoup(html.read(), "html.parser") institutions = [] # Iterate through each course's metadata container for course_meta in bsObj.find_all('div', class_='course-meta col-sm-12'): # Look for all <dt> tags in this container for dt in course_meta.find_all('dt'): # Check if this <dt> is the "Institution" label if dt.text.strip() == "Institution": # Grab the next sibling <dd> (the actual institution value) institution_dd = dt.find_next_sibling('dd') if institution_dd: # Make sure the <dd> exists to avoid errors institutions.append(institution_dd.text.strip()) break # No need to check other <dt>s once we find the right one # Print your results for inst in institutions: print(inst)
Option 2: More Concise Version
If you prefer a tighter approach, you can use a lambda to target the "Institution" <dt> directly:
from urllib.request import urlopen from bs4 import BeautifulSoup html = urlopen("http://www.oeconsortium.org/courses/search/?search=statistics") bsObj = BeautifulSoup(html.read(), "html.parser") institutions = [] for course_meta in bsObj.find_all('div', class_='course-meta col-sm-12'): # Find the <dt> with text containing "Institution" institution_label = course_meta.find('dt', string=lambda t: t and "Institution" in t.strip()) if institution_label: # Get the corresponding <dd> institution_value = institution_label.find_next_sibling('dd') if institution_value: institutions.append(institution_value.text.strip()) print(institutions)
Key Takeaways
- Always narrow down your scope first: start with each course's container (
div.course-meta col-sm-12) before digging into child elements. - Use
find_next_sibling()to pair<dt>labels with their corresponding<dd>values—this works perfectly because they're adjacent in the HTML structure. - Add checks (like
if institution_dd:) to handle cases where a course might not have an Institution listed, preventing AttributeErrors.
内容的提问来源于stack exchange,提问作者RichardS

