Python Selenium如何同时获取div内文本与a标签的href属性值
你可以在遍历每个class2元素时,基于当前元素上下文定位内部的a标签,调用get_attribute('href')方法即可获取链接属性,同时可以单独定位各文本节点取值,比直接拆分整体text更稳定:
完整代码示例
elements = driver.find_elements_by_xpath("//div[@class='class1']/div[@class='class2']") for e in elements: # 获取a标签的href属性 target_link = e.find_element_by_xpath(".//div[@class='class4']/a").get_attribute("href") # 分别获取各文本内容 text1 = e.find_element_by_xpath(".//div[@class='class3']").text.strip() text2 = e.find_element_by_xpath(".//a/h2").text.strip() text3 = e.find_element_by_xpath(".//p[@class='class5']").text.split(" ", 1)[0].strip() text4 = e.find_element_by_xpath(".//span[@class='class6']").text.strip() # 输出所有需要的内容 print("链接:", target_link) print("Text1:", text1) print("Text2:", text2) print("Text3:", text3) print("Text4:", text4) print("------------------------------------------")
关键说明
- 子元素定位的xpath开头加
.,代表仅在当前遍历到的e元素的后代节点中查找,不会全局匹配其他class1下的内容,避免取值混乱 - 如果不需要单独拆分Text3和Text4,直接取
e.find_element_by_xpath(".//p[@class='class5']").text就能同时拿到两者的内容 - 如果你的Selenium版本为4.x,原有的
find_element_by_xpath方法已废弃,建议改用以下写法:from selenium.webdriver.common.by import By # 定位父元素 elements = driver.find_elements(By.XPATH, "//div[@class='class1']/div[@class='class2']") for e in elements: target_link = e.find_element(By.XPATH, ".//div[@class='class4']/a").get_attribute("href") # 后续其他取值逻辑和上面一致
内容的提问来源于stack exchange,提问作者psmith
相关产品推荐
相关产品推荐

