如何用BeautifulSoup获取<li>标签?仅获取首个标签的问题排查
Fix: Only Getting First Tag with BeautifulSoup
Hey there! Let's troubleshoot why you're only pulling the first <li> tag when scraping that CCNA exam answers page. This is a super common issue, and it's usually easy to fix.
Most Likely Cause: Using find() Instead of find_all()
The #1 reason you're only getting one <li> is probably because you used page_soup.find("li") instead of page_soup.find_all("li").
find()returns only the first matching element in the HTML.find_all()returns a list of all matching elements—exactly what you need for extracting all list items.
Corrected Full Code
Here's your code updated to grab every <li> tag on the page:
from urllib.request import urlopen as uReq from bs4 import BeautifulSoup as soup url = 'https://ccnav6.com/ccna-4-chapter-1-exam-answers-2017-v5-0-3-v6-0-full-100.html' # Fetch and read the page content uClient = uReq(url) page_html = uClient.read() uClient.close() # Parse the HTML with BeautifulSoup page_soup = soup(page_html, "html.parser") # Get ALL <li> tags (not just the first one!) all_li_tags = page_soup.find_all("li") # Loop through and print each item to verify for li in all_li_tags: print(li.text.strip())
If You Still Have Issues: Target the Right Container
Sometimes, pages have multiple <ul>/<ol> sections, and you might want to only extract <li> tags from the specific Q&A list. To do this:
- Right-click the page → View Page Source (or Inspect Element)
- Find the parent container (like a
<ul>with a unique class or ID) that holds all the exam question/answer list items. - Update your code to target that container first:
# Example: If the Q&A list is in a <ul> with class 'exam-questions' qa_container = page_soup.find("ul", class_="exam-questions") if qa_container: all_qa_items = qa_container.find_all("li") for item in all_qa_items: print(item.text.strip()) else: print("Could not find the Q&A container. Check the page source for the correct selector.")
Quick Notes
- If the page uses JavaScript to load content dynamically (unlikely for this static exam page),
urllibwon't capture it. In that case, you'd need a tool like Selenium to render the page first. - Always close the
uClientconnection after reading the page to avoid resource leaks.
Give this a shot—you should see all the list items now!
内容的提问来源于stack exchange,提问作者Time4Boom
相关产品推荐
相关产品推荐

