You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Beautiful Soup爬取Instagram标题时遇NoneType TypeError求助

问题解决方法

你的代码存在几个关键问题,直接导致了NoneType错误和异常:

1. Instagram反爬拦截无标识请求

现在Instagram会拦截不带浏览器标识的请求,直接用requests.get(url)返回的是反爬页面,没有正常的<title>标签,所以doc.title会返回None,后续执行username in name时自然触发TypeError。

2. 函数外的无效代码

最后一行的print(name)完全多余,name是instagram函数内部的局部变量,这里直接调用会触发NameError。

3. 错误的title判断逻辑

即使doc.title存在,它是BeautifulSoup的Tag对象,不能直接用username in name判断,需要提取它的文本内容(name.string或转成字符串)。

修复后的代码

from bs4 import BeautifulSoup
import requests

def main():
    username = input("Enter username: ")
    instagram(username)

def instagram(username):
    url = f"https://www.instagram.com/{username}"
    # 添加浏览器请求头,模拟正常访问
    headers = {
        "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/114.0.0.0 Safari/537.36"
    }
    
    try:
        results = requests.get(url, headers=headers)
        # 先检查请求是否成功
        results.raise_for_status()
        
        doc = BeautifulSoup(results.text, "html.parser")
        title_tag = doc.title
        
        if title_tag:
            title_text = title_tag.string.strip()
            if username in title_text:
                print(f"找到用户页面标题:{title_text}")
            else:
                print("用户名不匹配,可能是重定向到其他页面")
        else:
            print("未找到页面标题,可能被反爬拦截或用户不存在")
    except requests.exceptions.RequestException as e:
        print(f"请求出错:{e}")

if __name__ == "__main__":
    main()

关键修改说明

  • 添加headers参数,带上浏览器的User-Agent,避免被Instagram反爬拦截
  • 加入try-except处理请求异常(比如网络错误、404/500状态码)
  • 先判断title_tag是否存在,再处理文本内容
  • 移除了函数外无效的print(name)代码
  • 使用f-string拼接URL更简洁

内容的提问来源于stack exchange,提问作者Hamza Khan

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.31 15:37:35