如何将特定XPath转为通用路径批量提取雅虎财经页面所有新闻链接
问题原因
你当前使用的XPath中li[1]的下标是固定选取列表的第一个元素,所以只能获取到单篇新闻。同时多层级嵌套的路径写法容错性极低,页面结构稍有调整就会匹配失效。
修正后代码
import time from selenium import webdriver # 若使用Selenium 4.0+版本,取消注释下一行导入 # from selenium.webdriver.common.by import By # 此处补充你的webdriver初始化逻辑,比如 driver = webdriver.Chrome() driver.maximize_window() driver.implicitly_wait(10) driver.get("https://finance.yahoo.com/quote/AAPL/news?p=AAPL") links = [] dates = [] # 滚动加载更多新闻 for i in range(20): driver.execute_script("window.scrollBy(0, 250)") time.sleep(1) # 匹配所有新闻列表项,移除固定下标[1] # Selenium 4.0+写法:all_items = driver.find_elements(By.XPATH, '//*[@id="latestQuoteNewsStream-0-Stream"]/ul/li') all_items = driver.find_elements_by_xpath('//*[@id="latestQuoteNewsStream-0-Stream"]/ul/li') for item in all_items: try: # 改用特征匹配,不写死层级,容错性更高 news_a = item.find_element_by_xpath('.//h3/a[@href]') news_date = item.find_element_by_xpath('.//div[contains(@class,"byline")]//span[contains(@class,"date")]') # 要新闻标题就取news_a.text,要链接就取news_a.get_attribute('href') links.append(news_a.get_attribute('href')) dates.append(news_date.text) except Exception: # 跳过广告位、空占位项等非新闻节点 continue driver.quit()
核心修改说明
- 移除XPath中
li后的[1]固定下标,匹配列表内所有新闻对应的li节点 - 放弃多层级嵌套的路径写法,改用特征标签模糊匹配,页面内部嵌套结构微调也不会失效
- 增加异常捕获逻辑,自动跳过列表中的广告、无效占位项
- 修正原代码中未导入
time库、导入未使用的依赖的问题 - 兼容Selenium新旧版本写法,可根据自己的环境选择对应调用方式
内容的提问来源于stack exchange,提问作者huy
相关产品推荐
相关产品推荐

