You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何处理HTML中多行多空格类名的元素定位问题?

问题:定位含多行空格的类名元素失败

我需要定位HTML中指定类名的所有元素,但从浏览器开发者工具复制的类名包含多行及大量空格,内容如下:

ll-sets-words__row
            false
        ```
尝试用Selenium和BeautifulSoup通过`CLASS_NAME`定位失败,改用`CSS_SELECTOR`或`XPATH`仅能找到单个元素,无法获取全部目标元素。以下是我的代码示例:

```python
from selenium import webdriver
from selenium.webdriver.chrome.service import Service
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time
import pickle
from bs4 import BeautifulSoup

url1 = 'https://lingualeo.com/'
url2 ='https://lingualeo.com/ru/dictionary/vocabulary/my'

s = Service('C:\\Users\\user\\Desktop\\chromedriver-win64\\chromedriver.exe')
options = webdriver.ChromeOptions()

options.add_argument('--excludeSwitches')
options.add_argument('--no-sandbox')
options.add_argument('--disable-dev-shm-usage')
options.add_argument('--disable-blink-features=AutomationControlled')

browser = webdriver.Chrome(service=s, options=options)
browser.maximize_window()
wait = WebDriverWait(browser, 600)
browser.get(url1)
time.sleep(7)

cookies = pickle.load((open('lingua_cookies.pkl', 'rb')))
for cookie in cookies:
    browser.add_cookie(cookie)
time.sleep(2)
browser.get(url2)

time.sleep(10)

class_name = """
                ll-sets-words__row
                false
            """

try:
    entire_dict = browser.find_element(By.CLASS_NAME, 'll-page-vocabulary__sets-words__table')
    print('it worked here1')
    words = entire_dict.find_elements(By.CLASS_NAME, class_name)
    for e in words:
        print(e.text)
except:
    print('error')
    browser.quit()
browser.quit()

代码在赋值words变量时出错,仍需通过类名定位所有目标元素。


解决方案

你复制的类名其实是两个独立的CSS类:ll-sets-words__row和false,多行和空格只是开发者工具的格式化显示,并不是类名的一部分。

修复方案:

1. 直接用核心类名定位

如果ll-sets-words__row已经能唯一标识目标元素,直接用这个类名即可:

words = entire_dict.find_elements(By.CLASS_NAME, 'll-sets-words__row')

2. 用CSS选择器匹配多类

如果需要同时匹配两个类(确保元素同时拥有这两个类),用CSS选择器的多类语法(类名用.连接,不要加空格):

words = entire_dict.find_elements(By.CSS_SELECTOR, '.ll-sets-words__row.false')

3. 自动清理复制的类名字符串

如果必须使用复制的原始字符串,先清理掉换行、空格,拆分出独立类名再组合:

# 清理类名字符串,拆分出有效类名
clean_classes = [cls.strip() for cls in class_name.split() if cls.strip()]
# 转换为CSS选择器格式
css_selector = '.'.join(clean_classes)
# 定位元素
words = entire_dict.find_elements(By.CSS_SELECTOR, css_selector)

失败原因说明:

  • By.CLASS_NAME只接受单个类名,不能传入包含空格或多行的字符串,否则会被当作一个完整的类名(而页面中不存在这样的类)。
  • 之前用CSS/XPATH只找到单个元素,大概率是写法错误——比如误加了空格(空格在CSS中代表后代选择器,不是多类匹配)。

内容的提问来源于stack exchange,提问作者mr.Jenkins

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.30 21:12:50