Python Selenium网页爬虫技术求助:如何获取目标页面最后一条帖子并提取其中完整链接
解决方案:修改Selenium爬虫获取最后一条帖子及完整链接
嘿,我来帮你调整这个爬虫脚本,实现只抓取页面最后一条帖子,同时提取里面的完整网络链接。咱们一步步来改:
核心修改点
1. 只获取最后一条评论
原来的代码会遍历所有评论,我们只需要取评论列表的最后一个元素就行,也就是comments[-1],这样就不用循环处理所有内容了。
2. 提取完整网络链接
之前你用.text只能拿到文本内容,要获取完整链接,得定位到评论里的<a>标签,然后提取它的href属性——这才是真正的完整链接地址。另外注意,一条评论里可能有多个链接,所以要遍历所有<a>标签来获取。
3. 替换已弃用的Selenium方法
你用的find_elements_by_class_name这类方法在Selenium 4+版本已经被弃用了,建议换成find_elements(By.CLASS_NAME, ...)这种新写法,避免后续版本兼容性问题。
修改后的完整代码
#!/usr/bin/env python3 from selenium import webdriver from selenium.webdriver.common.by import By from selenium.webdriver.chrome.options import Options from selenium.webdriver.support import expected_conditions as EC from selenium.webdriver.support.ui import WebDriverWait url = "https://www.dealabs.com/discussions/suivi-erreurs-de-prix-1063390?page=9999" options = Options() options.headless = True driver = webdriver.Chrome(options=options) driver.get(url) # 接受cookies(若后续页面更新,XPATH可能需要重新定位) button = WebDriverWait(driver, 2).until( EC.element_to_be_clickable((By.XPATH, "/html/body/main/div[4]/div[1]/div/div[1]/div[2]/button[2]/span")) ) button.click() # 获取所有评论,直接取最后一条 comments = driver.find_elements(By.CLASS_NAME, "commentList-item") if comments: # 先判断是否有评论,避免空列表报错 last_comment = comments[-1] # 提取帖子基础信息 _id = last_comment.get_attribute("id") author = last_comment.find_element(By.CLASS_NAME, 'userInfo-username').text content = last_comment.find_element(By.CLASS_NAME, 'userHtml-content').text timestamp = last_comment.find_element(By.CLASS_NAME, 'text--color-greyShade').text comment_url = f"{url}#{_id}" # 提取帖子里的完整网络链接 links = last_comment.find_elements(By.TAG_NAME, 'a') full_links = [link.get_attribute('href') for link in links if link.get_attribute('href')] # 打印结果 print("Posté par", author) print(content) print("Publication:", timestamp) print("Lien du commentaire:") print(comment_url) if full_links: print("\n帖子中的完整链接:") for link in full_links: print(link) print('-' * 30) else: print("页面中未找到任何评论") driver.close()
额外说明
- 如果页面的cookies按钮XPATH后续失效,你可以用浏览器开发者工具重新定位元素,换成更稳定的定位方式(比如
By.CSS_SELECTOR)。 - 代码里加了
if comments:的判断,防止页面没有评论时抛出索引错误,让脚本更健壮。
内容的提问来源于stack exchange,提问作者JohnDoe
相关产品推荐
相关产品推荐

