使用BeautifulSoup4网页爬取时无法获取目标链接的解决方案咨询
解决BeautifulSoup4爬取CNN Indonesia搜索结果链接失败的问题
问题分析
你的代码无法获取目标链接,大概率是以下几个原因:
- 缺少请求头模拟浏览器,服务器返回了非预期的页面内容(比如反爬拦截)
- 页面DOM结构发生变化,原选择器
div.media_rows无法定位到正确的结果区域 - 未对
results为空的情况做容错处理,会直接抛出异常中断程序
解决方案
针对上述问题,给出以下修复步骤:
添加请求头模拟浏览器
服务器会检测请求来源,添加User-Agent等头信息可以避免被拦截。更新页面选择器
当前CNN Indonesia搜索结果的文章列表区域是div.list.media_rows,每个文章项包含在div.list_item中,目标文章链接是该节点下的a标签(需过滤掉分页等无关链接)。添加容错处理
当页面加载异常或选择器无法定位元素时,跳过当前循环避免程序崩溃。
修改后的代码
import requests from bs4 import BeautifulSoup import json url_web = { "cnn" : "https://www.cnnindonesia.com/search/?query=citayam&page=", "detik" : "https://www.detik.com/search/searchall?query=citayam&siteid=2", "kompas" : "https://search.kompas.com/search/?q=citayam&submit=Submit" } list_cnn = [] # 添加请求头,模拟Chrome浏览器 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36" } for i in range(1, 33): URL = url_web['cnn'] + str(i) print(f"{i}/32 - {URL}") # 修正计数显示错误 try: page = requests.get(URL, headers=headers) page.raise_for_status() # 捕获HTTP请求错误 soup = BeautifulSoup(page.content, 'html.parser') # 定位正确的结果列表区域 results = soup.find("div", class_="list media_rows") if not results: print(f"第{i}页未找到结果区域,跳过") continue # 遍历每个文章项,提取目标链接 for item in results.find_all("div", class_="list_item"): a_tag = item.find("a") if a_tag and 'href' in a_tag.attrs: href = a_tag.get('href') # 确保是完整的文章链接,避免相对路径 if not href.startswith('http'): href = "https://www.cnnindonesia.com" + href list_cnn.append(href) print(href) print(f"当前已收集{len(list_cnn)}条链接") except Exception as e: print(f"第{i}页爬取失败: {str(e)}") continue root_path = 'gdrive/My Drive/analisa_cfw/' with open(root_path+'list_cnn.json', "w", encoding='utf8') as outfile: json.dump(list_cnn, outfile, ensure_ascii=False) print("Tokenized_sent json saved!")
关键改动说明
- 添加
headers模拟浏览器请求,避免被服务器拦截 - 修正计数显示错误(原代码中计数逻辑有误)
- 使用更精准的选择器定位文章项,过滤无关链接
- 添加
try-except和空值判断,增强程序稳定性 - 处理相对路径链接,确保保存的是完整URL
内容的提问来源于stack exchange,提问作者rhamss794
相关产品推荐
相关产品推荐

