You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用BeautifulSoup实现fb_pagZ li a选择器查询及href提取

How to Target fb_pagZ li a and Extract href with BeautifulSoup

Got it, let's fix this for you! You're already halfway there—since you know the CSS selector, BeautifulSoup has straightforward ways to implement that logic, either by using CSS selectors directly or chaining find()/find_all() calls. Here's how to do both:

1. Use select() (Direct CSS Selector Support)

BeautifulSoup's select() method natively accepts standard CSS selector syntax. Just remember to add a . before fb_pagZ (since it's a class selector in CSS), and you can grab all matching <a> tags in one step:

from bs4 import BeautifulSoup

# Replace with your actual HTML content
html = '''<div class="fb_pagZ"><li><a href="page1.html">Page 1</a></li><li><a href="page2.html">Page 2</a></li></div>'''

soup = BeautifulSoup(html, 'html.parser')

# Target all <a> tags inside <li> elements under the .fb_pagZ container
target_links = soup.select('.fb_pagZ li a')

# Extract the href attribute from each link
for link in target_links:
    href = link.get('href')
    print(href)  # Outputs: page1.html, page2.html

If you only need the first matching link, swap select() with select_one('.fb_pagZ li a').

2. Chain find()/find_all() (Build on Your Existing Code)

If you want to expand on the code you already have (finding the .fb_pagZ element first), you can chain subsequent searches to dig into its child nodes:

# First get the .fb_pagZ container (use find_all if there are multiple such elements)
pag_container = soup.find(class_='fb_pagZ')

if pag_container:  # Always check if the element exists to avoid errors
    # Find all <li> elements inside the container
    list_items = pag_container.find_all('li')
    for li in list_items:
        # Grab the <a> tag inside each <li>
        a_tag = li.find('a')
        if a_tag:
            print(a_tag.get('href'))

# Or a more concise chained version:
if pag_container:
    all_links = pag_container.find_all('a')  # Finds all <a> tags under the container
    for link in all_links:
        print(link.get('href'))

Quick Tip

Your original CSS selector fb_pagZ li a was missing the . prefix for the class—CSS requires .classname to target elements by class, which is why we use .fb_pagZ in the select() method.

内容的提问来源于stack exchange,提问作者John

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.26 09:10:18