You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python爬虫:处理含空格的class属性值时出现报错问题

Fixing Class Attribute Matching with Spaces in BeautifulSoup

Hey there! Let's break down why your code is throwing errors and how to fix it properly.

The Root Problem

First, let's clarify what that HTML class attribute means:

<div class="row No 4">

In HTML, spaces in the class attribute don't belong to a single class name—they separate multiple distinct classes. So this <div> actually has three classes: row, No, and 4. Your original approaches were treating it as a single class name, which is why they failed.

Why Your Original Code Failed

  1. First attempt:

    files = bsObj.find_all("div", {"class":"row No 4"}) for file in files: print (file.get_text())
    

    This has two issues:

    • You're trying to assign a variable and run a for loop in one line without proper syntax (needs a newline or semicolon to separate the two statements).
    • Passing "row No 4" as the class value tells BeautifulSoup to look for elements where the class attribute is exactly that string—which doesn't exist, since the element's class is split into three separate values.
  2. Second attempt:

    files = bsObj.find_all("div", {"class":"row.No.4"}) for file in files: print (file.get_text())
    

    Here you're still trying to match a single class name string ("row.No.4"), but the element doesn't have a class with dots in it. The dot notation works in CSS selectors, not when directly specifying the class attribute value.

Correct Solutions

Solution 1: Use a List of Class Names

BeautifulSoup lets you pass a list of class names to find_all()—it will match elements that have all of those classes:

# Note: Use class_ instead of class (since class is a reserved Python keyword)
files = bsObj.find_all("div", class_=["row", "No", "4"])
for file in files:
    print(file.get_text())

Solution 2: Use CSS Selectors with select()

For CSS selectors, multiple class selectors (each prefixed with a dot) target elements that have all those classes. Use the select() method instead of find_all():

files = bsObj.select("div.row.No.4")
for file in files:
    print(file.get_text())

This works because .row.No.4 translates to "find <div> elements that have the classes row, No, and 4 all at once".

内容的提问来源于stack exchange,提问作者Kaj Kam

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 09:49:37