You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup获取<li>标签?仅获取首个标签的问题排查

Fix: Only Getting First
  • Tag with BeautifulSoup
  • Hey there! Let's troubleshoot why you're only pulling the first <li> tag when scraping that CCNA exam answers page. This is a super common issue, and it's usually easy to fix.

    Most Likely Cause: Using find() Instead of find_all()

    The #1 reason you're only getting one <li> is probably because you used page_soup.find("li") instead of page_soup.find_all("li").

    • find() returns only the first matching element in the HTML.
    • find_all() returns a list of all matching elements—exactly what you need for extracting all list items.

    Corrected Full Code

    Here's your code updated to grab every <li> tag on the page:

    from urllib.request import urlopen as uReq
    from bs4 import BeautifulSoup as soup
    
    url = 'https://ccnav6.com/ccna-4-chapter-1-exam-answers-2017-v5-0-3-v6-0-full-100.html'
    
    # Fetch and read the page content
    uClient = uReq(url)
    page_html = uClient.read()
    uClient.close()
    
    # Parse the HTML with BeautifulSoup
    page_soup = soup(page_html, "html.parser")
    
    # Get ALL <li> tags (not just the first one!)
    all_li_tags = page_soup.find_all("li")
    
    # Loop through and print each item to verify
    for li in all_li_tags:
        print(li.text.strip())
    

    If You Still Have Issues: Target the Right Container

    Sometimes, pages have multiple <ul>/<ol> sections, and you might want to only extract <li> tags from the specific Q&A list. To do this:

    1. Right-click the page → View Page Source (or Inspect Element)
    2. Find the parent container (like a <ul> with a unique class or ID) that holds all the exam question/answer list items.
    3. Update your code to target that container first:
    # Example: If the Q&A list is in a <ul> with class 'exam-questions'
    qa_container = page_soup.find("ul", class_="exam-questions")
    if qa_container:
        all_qa_items = qa_container.find_all("li")
        for item in all_qa_items:
            print(item.text.strip())
    else:
        print("Could not find the Q&A container. Check the page source for the correct selector.")
    

    Quick Notes

    • If the page uses JavaScript to load content dynamically (unlikely for this static exam page), urllib won't capture it. In that case, you'd need a tool like Selenium to render the page first.
    • Always close the uClient connection after reading the page to avoid resource leaks.

    Give this a shot—you should see all the list items now!

    内容的提问来源于stack exchange,提问作者Time4Boom

    相关产品推荐
    方舟 Agent Plan

    超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

    最近更新时间:2026.05.21 06:23:35