如何用Regex或Beautiful Soup抓取Instagram简介中的隐藏外部链接
如何抓取Instagram用户主页简介中的隐藏外部网址
嘿,我刚好踩过这个坑!Instagram确实把主页简介里的外部链接藏在页面的JavaScript代码块里,没法直接用Beautiful Soup去抓常规的<a>标签,不过用正则表达式或者结合Beautiful Soup定位JS代码再解析,就能轻松搞定。给你两种亲测有效的方案:
方案一:直接用正则表达式匹配
这种方法简单直接,因为目标链接始终出现在"external_url":"..."这个键值对里,我们可以直接用正则匹配这个模式:
import requests import re # 替换成你要抓取的用户主页URL target_url = "https://www.instagram.com/brittanyannecohen/" # 模拟浏览器请求头,避免被Instagram反爬拦截 headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: response = requests.get(target_url, headers=headers) response.raise_for_status() # 抛出请求异常 # 匹配external_url对应的链接,注意转义双引号 url_pattern = re.compile(r'"external_url":"(.*?)"') match_result = url_pattern.search(response.text) if match_result: # 去除链接前后的空格 external_link = match_result.group(1).strip() print(f"抓取到的外部链接:{external_link}") else: print("该用户主页没有设置外部网址") except Exception as e: print(f"请求或抓取失败:{str(e)}")
方案一注意点
- 一定要设置真实的User-Agent,不然很容易被Instagram返回403禁止访问
- 如果遇到匹配失败,可以检查正则表达式是否需要调整(比如Instagram可能在键值对前后加了额外空格,不过上面的正则已经兼容这种情况)
方案二:解析Instagram的共享数据(更稳定)
Instagram会把用户的所有公开数据存在window._sharedData这个JS对象里,我们可以定位到这个脚本块,把它转换成Python字典后再提取链接,这种方法更不容易因为页面小改动失效:
import requests from bs4 import BeautifulSoup import re import json target_url = "https://www.instagram.com/brittanyannecohen/" headers = { "User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36" } try: response = requests.get(target_url, headers=headers) response.raise_for_status() soup = BeautifulSoup(response.text, "html.parser") # 找到包含_sharedData的script标签 target_script = None for script in soup.find_all("script", type="text/javascript"): if script.string and "window._sharedData" in script.string: target_script = script.string break if target_script: # 提取出JSON格式的数据部分 data_pattern = re.compile(r'window._sharedData = (.*?);') json_match = data_pattern.search(target_script) if json_match: user_data = json.loads(json_match.group(1)) # 导航到用户数据中的external_url字段 external_link = user_data["entry_data"]["ProfilePage"][0]["graphql"]["user"]["external_url"] if external_link: print(f"抓取到的外部链接:{external_link.strip()}") else: print("该用户主页没有设置外部网址") else: print("未能提取到共享数据") else: print("未找到包含用户数据的脚本块") except Exception as e: print(f"请求或抓取失败:{str(e)}")
方案二优势
- 直接解析官方的数据结构,比正则匹配更稳定,就算页面JS代码有小的格式调整,只要数据结构不变就能正常抓取
- 除了外部链接,还能顺便获取用户的粉丝数、帖子数等其他公开数据
通用注意事项
- 频繁请求Instagram可能会触发反爬机制,建议每次请求后加1-2秒的延迟,或者使用代理IP
- 如果用户设置了隐私账号,未登录状态下无法获取到这些数据,需要先处理登录逻辑
- 有些用户可能没有添加外部网址,一定要做判空处理,避免代码报错
内容的提问来源于stack exchange,提问作者Bob
相关产品推荐
相关产品推荐

