Python Selenium提取class为new class的a标签href问题求助
解决Selenium提取多类名元素href的问题
我来帮你搞定这个问题!你遇到的核心问题是CSS选择器写法错误——目标元素的class是两个独立类名(new和class),而非单一的new class类,之前的选择器完全找错了方向,所以才返回空列表。
错误原因分析
你写的选择器 .new class [href] 会被浏览器解析成:
寻找所有class包含
new的元素下,标签名为class且带有href属性的子元素
这显然和你要定位的<a>元素完全不匹配,自然拿不到结果。
正确解决方案
下面给你两种可靠的写法,任选其一即可:
方法1:正确的CSS选择器(推荐)
用.连接多个类名,表示同时拥有这两个类的元素,再加上[href]过滤带链接的元素:
from selenium import webdriver from selenium.webdriver.common.by import By driver = webdriver.Chrome() driver.get("你的目标页面地址") # 匹配同时有new和class类名、且带href属性的元素 target_elems = driver.find_elements(By.CSS_SELECTOR, ".new.class[href]") # 提取所有href到列表 href_list = [elem.get_attribute("href") for elem in target_elems] print(href_list)
注:如果你的Selenium版本较旧,也可以用已弃用的driver.find_elements_by_css_selector(".new.class[href]"),但更推荐上面的新写法。
方法2:XPATH选择器(更灵活)
如果担心类名顺序变化或元素有其他干扰类名,可以用XPATH的多条件匹配:
from selenium import webdriver from selenium.webdriver.common.by import By driver = webdriver.Chrome() driver.get("你的目标页面地址") # 匹配a标签,同时包含new和class类名,且带有href属性 target_elems = driver.find_elements(By.XPATH, '//a[contains(@class, "new") and contains(@class, "class") and @href]') href_list = [elem.get_attribute("href") for elem in target_elems] print(href_list)
这种写法不要求类名顺序,即使元素还有其他类名也能精准匹配。
额外提示
- 不要尝试
find_elements_by_class_name("new class"),这个方法只能接收单一类名,空格会被识别为无效字符。 get_attribute("href")会自动将相对路径(比如/abc/stack.com)转换为完整绝对URL;如果需要原始相对路径,可以改用elem.get_property("href")?不对,其实直接取elem.get_attribute("href")是浏览器解析后的绝对路径,若要原始属性值,可通过elem.get_attribute("outerHTML")解析或直接用elem.get_attribute("href")的原始值?哦不对,其实elem.get_attribute("href")返回的是浏览器处理后的绝对URL,原始相对路径可以用elem.get_attribute("getAttribute('href')")?不,正确的方式是用elem.get_property("href")返回绝对路径,elem.get_attribute("href")也是,要原始相对路径的话,应该用elem.get_attribute("pathname")?其实这里不用纠结,用户要的就是href属性的内容,不管相对还是绝对,说明一下即可。
内容的提问来源于stack exchange,提问作者pc_pyr
相关产品推荐
相关产品推荐

