You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Python中BeautifulSoup调用find_all返回空列表问题求助

解决IP信息爬取返回空列表的问题

核心问题分析及解决步骤

1. 被网站反爬机制拦截

大部分网站会拒绝无请求头的爬虫请求,导致返回的内容并非正常页面,自然无法定位目标元素。

修改代码添加模拟浏览器的请求头:

import requests as r
from bs4 import BeautifulSoup

def ip_infos():
    url = "https://whatismyipaddress.com/ip/"
    ip = input("Adresse IP à rechercher : ")
    url_completed = url + ip
    
    # 加入请求头模拟浏览器访问
    headers = {
        'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36'
    }
    y = r.get(url_completed, headers=headers)
    
    # 先验证请求是否成功
    if y.status_code != 200:
        print(f"请求失败,状态码: {y.status_code}")
        return
    
    ip_check_soup = BeautifulSoup(y.content, features="html5lib")
    # 页面结构已更新,调整选择器匹配当前页面元素
    details_container = ip_check_soup.find('div', class_='ip-details')
    if not details_container:
        print("未找到IP信息区域")
        return
    
    # 提取具体信息项
    info_rows = details_container.find_all('div', class_='row')
    for row in info_rows:
        label = row.find('div', class_='col-md-3').text.strip()
        value = row.find('div', class_='col-md-9').text.strip()
        print(f"{label}: {value}")

ip_infos()

2. 目标页面结构已变更

你原代码使用的find_all('p', class_='information')选择器对应的页面元素已被网站修改。当前该网站的IP详情页面,信息存储在div.ip-details下的div.row结构中,必须调整选择器才能正确抓取内容。

额外注意事项

  • 频繁请求可能触发反爬限制,建议添加请求间隔(如time.sleep(2))
  • 若部分信息需JavaScript渲染才能加载,可尝试使用selenium模拟完整浏览器加载流程

内容的提问来源于stack exchange,提问作者Blackax

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 02:51:07