You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python3.2中用soup.findAll获取class值及多class匹配问题排查

Let's tackle your two BeautifulSoup questions one by one—these are common snags when working with HTML parsing, so let's break them down clearly:

1. How to get an element's class attribute value with soup.findAll() in Python 3.2

First, make sure you're using BeautifulSoup 4 (bs4)—it’s fully compatible with Python 3.2 and the standard for this kind of work. Here’s how to do it:

  • Use findAll() (or its alias find_all()—they work identically) to locate elements by class. Note we use class_ (with an underscore) instead of plain class because class is a reserved keyword in Python.
  • Once you’ve fetched the elements, access the class attribute like a dictionary key: elem['class']. This returns a list of strings, since HTML elements can have multiple classes.

Example code to illustrate:

from bs4 import BeautifulSoup

# Sample HTML content
html = '''
<div class="basic-class">Test Element 1</div>
<div class="multi class group">Test Element 2</div>
'''

# Initialize the soup object
soup = BeautifulSoup(html, 'html.parser')

# Find all divs with "basic-class"
basic_elements = soup.findAll('div', class_='basic-class')
for elem in basic_elements:
    print(f"Class list: {elem['class']}")  # Output: ['basic-class']

# Find all divs containing "class" in their classes
multi_elements = soup.findAll('div', class_='class')
for elem in multi_elements:
    print(f"Class list: {elem['class']}")  # Output: ['multi', 'class', 'group']
2. Fixing multi-class element matching

Your custom match_class() approach failed because BeautifulSoup parses the class attribute into a list of individual class names, not a single space-separated string. Comparing against a full string like "col-lg-3 col-md-4 col-sm-6 bt-product-list" won’t match since the element’s class value is stored as a list.

You have two simple, reliable fixes:

Option 1: Use BeautifulSoup's built-in multi-class support

BeautifulSoup lets you pass a list of class names to class_, and it will match elements that have all those classes (order doesn’t matter):

# Target classes as a list
target_classes = ['col-lg-3', 'col-md-4', 'col-sm-6', 'bt-product-list']
# Find divs with all these classes
elements = soup.findAll('div', class_=target_classes)

Option 2: Fix your custom match_class function

If you prefer using a custom function, update it to check that every target class exists in the element’s class list:

def match_class(target_classes):
    def check_element(elem):
        # Get the element's class list (empty list if no class attribute exists)
        elem_classes = elem.get('class', [])
        # Return True only if all target classes are present
        return all(cls in elem_classes for cls in target_classes)
    return check_element

# Now use the function correctly
elements = soup.findAll(match_class(['col-lg-3', 'col-md-4', 'col-sm-6', 'bt-product-list']))

Either method will correctly locate the elements you’re targeting.

内容的提问来源于stack exchange,提问作者fatra nitha

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 06:57:56