You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

网页爬取SCMP新闻文章时无法获取正文内容求助

问题描述

我正在从SCMP网站爬取新闻文章,能够获取文章的标题与作者名称,但无法获取正文内容。尝试了两种方法均未成功,代码如下:

第一种方法

options = webdriver.ChromeOptions()

lists = ['disable-popup-blocking']

caps = DesiredCapabilities().CHROME
caps["pageLoadStrategy"] = "normal"

driver.get('https://www.scmp.com/news/asia/east-asia/article/3199400/japan-asean-hold-summit-tokyo-around-december-2023-japanese-official')
driver.implicitly_wait(5)

bsObj = BeautifulSoup(driver.page_source, 'html.parser')
text_res = bsObj.select('div[class="details__body body"]') 
    
text = ""
for item in text_res:
    if item.get_text() == "":
        continue
    text = text + item.get_text().strip() + "\n"   

第二种方法

options = webdriver.ChromeOptions()

driver = webdriver.Chrome(executable_path= r"E:\chromedriver\chromedriver.exe", options=options) #add your chrome path    

driver.get('https://www.scmp.com/news/asia/east-asia/article/3199400/japan-asean-hold-summit-tokyo-around-december-2023-japanese-official')
driver.implicitly_wait(5)

a = driver.find_element_by_class_name("details__body body").text
print(a)
问题分析
  1. 语法错误:第二种方法中find_element_by_class_name仅支持单个类名,不能用空格传递多个类名(details__body body是两个类),直接导致元素查找失败。
  2. 等待策略不足:implicitly_wait(5)的固定等待时间可能不足以让动态加载的正文完全渲染,尤其是网站有反爬延迟加载的情况。
  3. 选择器与内容结构不匹配:即使定位到父容器,正文可能嵌套在更下层的子元素(如<p>标签)中,直接提取父容器文本可能无法获取有效内容;另外网站可能更新了页面结构,原选择器已失效。
  4. 反爬拦截:SCMP可能通过浏览器特征检测、Cookie验证等方式限制爬虫,未处理的情况下页面可能无法正常加载正文。
解决方案

1. 修复元素定位语法

将多类选择改为css_selector,这是支持多类组合定位的正确方式:

a = driver.find_element_by_css_selector("div.details__body.body").text
print(a)

2. 改用显式等待确保内容加载

使用WebDriverWait等待正文元素可见或包含文本,比隐式等待更可靠:

from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.common.by import By

# 初始化driver代码...

driver.get('目标文章URL')
# 最长等待10秒,直到正文容器可见
wait = WebDriverWait(driver, 10)
body_container = wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, "div.details__body.body")))
# 提取容器内所有p标签的文本(SCMP正文通常嵌套在p标签中)
paragraphs = body_container.find_elements_by_tag_name("p")
full_text = "\n".join([p.text.strip() for p in paragraphs if p.text.strip()])
print(full_text)

3. 添加反爬规避参数

在ChromeOptions中加入模拟正常浏览器的配置,避免被识别为爬虫:

options = webdriver.ChromeOptions()
# 禁用自动化识别特征
options.add_argument("--disable-blink-features=AutomationControlled")
# 设置正常用户代理
options.add_argument("user-agent=Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36")
# 移除自动化提示
options.add_experimental_option("excludeSwitches", ["enable-automation"])
options.add_experimental_option('useAutomationExtension', False)

4. 验证页面结构

右键目标文章页面的正文区域,检查元素确认当前的类名和层级结构。如果SCMP更新了页面布局,需要同步调整选择器(例如正文可能迁移到article-body类容器下)。

内容的提问来源于stack exchange,提问作者Starlord22

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.12 20:05:32