You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Regex或Beautiful Soup抓取Instagram简介中的隐藏外部链接

如何抓取Instagram用户主页简介中的隐藏外部网址

嘿,我刚好踩过这个坑!Instagram确实把主页简介里的外部链接藏在页面的JavaScript代码块里,没法直接用Beautiful Soup去抓常规的<a>标签,不过用正则表达式或者结合Beautiful Soup定位JS代码再解析,就能轻松搞定。给你两种亲测有效的方案:

方案一:直接用正则表达式匹配

这种方法简单直接,因为目标链接始终出现在"external_url":"..."这个键值对里,我们可以直接用正则匹配这个模式:

import requests
import re

# 替换成你要抓取的用户主页URL
target_url = "https://www.instagram.com/brittanyannecohen/"
# 模拟浏览器请求头,避免被Instagram反爬拦截
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

try:
    response = requests.get(target_url, headers=headers)
    response.raise_for_status()  # 抛出请求异常
    
    # 匹配external_url对应的链接,注意转义双引号
    url_pattern = re.compile(r'"external_url":"(.*?)"')
    match_result = url_pattern.search(response.text)
    
    if match_result:
        # 去除链接前后的空格
        external_link = match_result.group(1).strip()
        print(f"抓取到的外部链接:{external_link}")
    else:
        print("该用户主页没有设置外部网址")
except Exception as e:
    print(f"请求或抓取失败:{str(e)}")

方案一注意点

  • 一定要设置真实的User-Agent,不然很容易被Instagram返回403禁止访问
  • 如果遇到匹配失败,可以检查正则表达式是否需要调整(比如Instagram可能在键值对前后加了额外空格,不过上面的正则已经兼容这种情况)

方案二:解析Instagram的共享数据(更稳定)

Instagram会把用户的所有公开数据存在window._sharedData这个JS对象里,我们可以定位到这个脚本块,把它转换成Python字典后再提取链接,这种方法更不容易因为页面小改动失效:

import requests
from bs4 import BeautifulSoup
import re
import json

target_url = "https://www.instagram.com/brittanyannecohen/"
headers = {
    "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36"
}

try:
    response = requests.get(target_url, headers=headers)
    response.raise_for_status()
    soup = BeautifulSoup(response.text, "html.parser")
    
    # 找到包含_sharedData的script标签
    target_script = None
    for script in soup.find_all("script", type="text/javascript"):
        if script.string and "window._sharedData" in script.string:
            target_script = script.string
            break
    
    if target_script:
        # 提取出JSON格式的数据部分
        data_pattern = re.compile(r'window._sharedData = (.*?);')
        json_match = data_pattern.search(target_script)
        if json_match:
            user_data = json.loads(json_match.group(1))
            # 导航到用户数据中的external_url字段
            external_link = user_data["entry_data"]["ProfilePage"][0]["graphql"]["user"]["external_url"]
            if external_link:
                print(f"抓取到的外部链接:{external_link.strip()}")
            else:
                print("该用户主页没有设置外部网址")
        else:
            print("未能提取到共享数据")
    else:
        print("未找到包含用户数据的脚本块")
except Exception as e:
    print(f"请求或抓取失败:{str(e)}")

方案二优势

  • 直接解析官方的数据结构,比正则匹配更稳定,就算页面JS代码有小的格式调整,只要数据结构不变就能正常抓取
  • 除了外部链接,还能顺便获取用户的粉丝数、帖子数等其他公开数据

通用注意事项

  • 频繁请求Instagram可能会触发反爬机制,建议每次请求后加1-2秒的延迟,或者使用代理IP
  • 如果用户设置了隐私账号,未登录状态下无法获取到这些数据,需要先处理登录逻辑
  • 有些用户可能没有添加外部网址,一定要做判空处理,避免代码报错

内容的提问来源于stack exchange,提问作者Bob

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.27 03:32:56