You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用BeautifulSoup4网页爬取时无法获取目标链接的解决方案咨询

解决BeautifulSoup4爬取CNN Indonesia搜索结果链接失败的问题

问题分析

你的代码无法获取目标链接,大概率是以下几个原因:

  • 缺少请求头模拟浏览器,服务器返回了非预期的页面内容(比如反爬拦截)
  • 页面DOM结构发生变化,原选择器div.media_rows无法定位到正确的结果区域
  • 未对results为空的情况做容错处理,会直接抛出异常中断程序

解决方案

针对上述问题,给出以下修复步骤:

  1. 添加请求头模拟浏览器
    服务器会检测请求来源,添加User-Agent等头信息可以避免被拦截。

  2. 更新页面选择器
    当前CNN Indonesia搜索结果的文章列表区域是div.list.media_rows,每个文章项包含在div.list_item中,目标文章链接是该节点下的a标签(需过滤掉分页等无关链接)。

  3. 添加容错处理
    当页面加载异常或选择器无法定位元素时,跳过当前循环避免程序崩溃。

修改后的代码

import requests
from bs4 import BeautifulSoup
import json

url_web = {
    "cnn" : "https://www.cnnindonesia.com/search/?query=citayam&page=",
    "detik" : "https://www.detik.com/search/searchall?query=citayam&siteid=2",
    "kompas" : "https://search.kompas.com/search/?q=citayam&submit=Submit"
}

list_cnn = []
# 添加请求头,模拟Chrome浏览器
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
}

for i in range(1, 33):
    URL = url_web['cnn'] + str(i)
    print(f"{i}/32 - {URL}")  # 修正计数显示错误
    try:
        page = requests.get(URL, headers=headers)
        page.raise_for_status()  # 捕获HTTP请求错误
        soup = BeautifulSoup(page.content, 'html.parser')
        # 定位正确的结果列表区域
        results = soup.find("div", class_="list media_rows")
        if not results:
            print(f"第{i}页未找到结果区域,跳过")
            continue
        # 遍历每个文章项,提取目标链接
        for item in results.find_all("div", class_="list_item"):
            a_tag = item.find("a")
            if a_tag and 'href' in a_tag.attrs:
                href = a_tag.get('href')
                # 确保是完整的文章链接,避免相对路径
                if not href.startswith('http'):
                    href = "https://www.cnnindonesia.com" + href
                list_cnn.append(href)
                print(href)
        print(f"当前已收集{len(list_cnn)}条链接")
    except Exception as e:
        print(f"第{i}页爬取失败: {str(e)}")
        continue

root_path = 'gdrive/My Drive/analisa_cfw/'
with open(root_path+'list_cnn.json', "w", encoding='utf8') as outfile:
    json.dump(list_cnn, outfile, ensure_ascii=False)
print("Tokenized_sent json saved!")

关键改动说明

  • 添加headers模拟浏览器请求,避免被服务器拦截
  • 修正计数显示错误(原代码中计数逻辑有误)
  • 使用更精准的选择器定位文章项,过滤无关链接
  • 添加try-except和空值判断,增强程序稳定性
  • 处理相对路径链接,确保保存的是完整URL

内容的提问来源于stack exchange,提问作者rhamss794

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.24 15:39:15