如何使用Selenium抓取经济时报存档页面表格内的新闻链接?
问题排查与修正方案
现有代码的核心错误点
- 链接属性归属错误:你要抓取的新闻链接是放在
<td>标签内部的<a>标签上,直接对<td>调用get_attribute('href')当然拿不到值 - 元素获取方法误用:
find_element()仅返回匹配到的第一个元素,要遍历表格所有行、所有单元格,需要使用返回元素列表的find_elements()方法 - 实例初始化冗余:你配置了禁用通知的Chrome实例完全没有被使用,实际发起请求的是无配置的默认实例,可能被页面通知弹窗拦截影响元素定位
修正后可运行代码
from selenium import webdriver from selenium.webdriver.chrome.options import Options from selenium.webdriver.chrome.service import Service from selenium.webdriver.common.by import By # 替换为你的chromedriver本地路径 driverLocation = "你的chromedriver存储路径" options = Options() options.add_argument("--disable-notifications") # 高版本selenium需通过Service传入驱动路径 s = Service(driverLocation) browser = webdriver.Chrome(service=s, options=options) url = 'https://economictimes.indiatimes.com/archive/year-2021,month-1.cms' browser.get(url) # 定位日历表格 table = browser.find_element(By.ID,"calender") # 获取表格所有行 rows = table.find_elements(By.TAG_NAME, "tr") # 遍历所有行和单元格提取链接 for row in rows: tds = row.find_elements(By.TAG_NAME, "td") for td in tds: # 跳过无链接的空白单元格 try: a_tag = td.find_element(By.TAG_NAME, "a") print(f"日期:{a_tag.text},对应新闻链接:{a_tag.get_attribute('href')}") except: continue # 抓取完成后关闭浏览器 browser.quit()
补充说明
如果仅需要获取指定单个日期的链接,也可以直接写xpath定位到对应a标签再获取href,不需要逐层遍历表格。
内容的提问来源于stack exchange,提问作者Swarnalakshmi Umamaheshwaran
相关产品推荐
相关产品推荐

