You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python Selenium遍历HTML表格时仅获取第一行问题求助

解决新浪热点新闻标题提取的遍历问题

嘿,咱们来捋捋为啥你爬新浪热点新闻时每次只拿到第一个标题——我之前处理表格布局的爬取也遇到过类似问题,下面说说可能的原因和解决办法:

常见问题根源

  • 误用单元素查找方法:比如用BeautifulSoup的find()而不是find_all(),前者只会返回匹配到的第一个元素,后者才会返回所有符合条件的结果,这是最常见的坑。
  • 遍历范围没限定:如果循环里每次都从页面根节点查找行元素,而不是基于当前板块的表格来定位,就会一直重复拿到第一个板块的第一行内容。
  • 页面动态加载:新浪部分热点内容可能是通过JS动态渲染的,用requests这类静态请求工具只能拿到初始HTML,看不到后续加载的新闻。

针对性解决方案

假设你用的是BeautifulSoup,这里给你一个正确的遍历逻辑示例:

1. 静态内容爬取(页面内容是静态渲染的情况)

import requests
from bs4 import BeautifulSoup

url = "http://news.sina.com.cn/hotnews/"
response = requests.get(url)
response.encoding = "utf-8"  # 确保编码正确,避免乱码
soup = BeautifulSoup(response.text, "html.parser")

# 遍历你感兴趣的板块(请根据实际HTML结构调整标签和类名)
for section in soup.find_all("div", class_="hotnews_section"):
    # 重点:在当前板块下查找所有表格行,而不是全局查找
    news_rows = section.find_all("tr")
    print(f"当前板块共找到 {len(news_rows)} 条新闻")
    
    # 遍历每一行提取标题
    for row in news_rows:
        title_link = row.find("a")
        if title_link:
            news_title = title_link.get_text(strip=True)
            print(news_title)

2. 动态内容处理(静态爬取不到完整内容的情况)

如果页面是JS动态加载的,就得用selenium模拟浏览器来获取完整内容:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

# 初始化浏览器(需要对应浏览器的驱动,比如ChromeDriver)
driver = webdriver.Chrome()
driver.get("http://news.sina.com.cn/hotnews/")

# 等待板块加载完成,避免因未加载完导致的空结果
wait = WebDriverWait(driver, 10)
sections = wait.until(EC.presence_of_all_elements_located((By.CLASS_NAME, "hotnews_section")))

for section in sections:
    # 限定在当前板块下查找所有表格行
    news_rows = section.find_elements(By.TAG_NAME, "tr")
    print(f"当前板块共找到 {len(news_rows)} 条新闻")
    
    for row in news_rows:
        try:
            title_link = row.find_element(By.TAG_NAME, "a")
            print(title_link.text.strip())
        except:
            # 跳过没有标题的行,避免报错中断程序
            continue

driver.quit()

几个关键注意点

  • 严格限定查找范围:必须在当前板块的节点下调用find_all(),而不是直接从soup全局查找,这是避免重复拿到第一个元素的核心。
  • 核对HTML结构:打开浏览器开发者工具(F12),确认板块、表格行、标题的实际标签和类名,不要凭猜测写选择器。
  • 容错处理:加入判断或异常捕获,避免因为某些行没有标题元素导致程序崩溃。

内容的提问来源于stack exchange,提问作者Sebastian

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.19 08:30:40