You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

求助:News24新闻文章内容爬取不全问题排查(附代码)

问题分析与解决建议

针对News24网站文章内容爬取不完整的问题,常见原因及对应解决方法如下:

1. 动态加载内容未处理

News24部分文章内容通过JavaScript动态渲染(如滚动加载剩余段落、点击展开全文),静态HTML解析无法获取完整内容。

解决方法:

使用浏览器自动化工具模拟页面渲染,等待所有内容加载完成后提取:

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
import time

driver = webdriver.Chrome()
driver.get(article_url)

# 等待核心内容容器加载完成(根据实际DOM结构调整选择器)
WebDriverWait(driver, 10).until(
    EC.presence_of_element_located((By.CLASS_NAME, "article-body"))
)

# 处理"阅读更多"按钮(若存在)
try:
    read_more_btn = driver.find_element(By.CLASS_NAME, "read-more")
    if read_more_btn.is_displayed():
        read_more_btn.click()
        time.sleep(1)
except:
    pass

# 提取所有段落内容
content_elements = driver.find_elements(By.CSS_SELECTOR, "div.article-body p")
full_content = "\n".join([elem.text.strip() for elem in content_elements])

driver.quit()

2. 内容选择器范围过窄

若仅使用单一选择器(如只抓取第一段),会导致内容截断。需确保选择器覆盖所有内容节点。

解决方法:

通过浏览器开发者工具定位完整内容的DOM结构,调整选择器:

from bs4 import BeautifulSoup

# 假设已获取页面HTML
soup = BeautifulSoup(html_content, "html.parser")
# 选择文章主体下的所有段落标签
content_paragraphs = soup.select("div.article-body > p")
full_content = "\n".join([p.get_text(strip=True) for p in content_paragraphs])

3. 反爬机制导致响应截断

News24可能对频繁无标识请求进行限制,返回不完整的页面内容。

解决方法:

  • 添加模拟浏览器的请求头:
import requests

headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36",
    "Accept-Language": "en-US,en;q=0.9"
}
response = requests.get(article_url, headers=headers)
response.encoding = response.apparent_encoding  # 确保编码正确
  • 增加请求间隔:
import time
time.sleep(2)  # 每次请求后暂停2秒,避免触发反爬

内容的提问来源于stack exchange,提问作者TG_

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.09 21:10:25