You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用Jsoup无法获取Instagram的video标签,求解决方法

Instagram视频下载问题:无法通过Jsoup定位视频资源

我是后端开发者,已有15年未接触HTML文档解析,对相关操作不太熟悉;同时我对Instagram的运作机制也处于学习阶段。

我尝试下载Instagram上的视频,按常规逻辑视频应该在<video>标签中,但无论通过何种方式遍历org.jsoup.nodes.Document的子元素,都无法识别到该标签。我试过调用Document.children().select(*)方法,怀疑Instagram隐藏了视频源,但毫无头绪。

我原本预期页面会存在og:video元标签(页面中title、img等元标签均正常存在),并尝试通过代码page.select("meta[property=og:video]").first().attr("content");获取视频地址,但该元标签并不存在。我还在InstagramDownloader类中编写了两个递归方法遍历所有节点和元素(方法来自另一Stack Overflow问题),但依然没找到获取视频的线索。我甚至不确定即便拿到视频的src URL,是否能成功完成下载。

相关代码如下:

public class Application {

    public static void main(String[] args) {

        try {
            login();
        } catch (Exception e) {
            e.printStackTrace();
        }
    }

    public static void login() throws IGLoginException, InterruptedException, ExecutionException{

        IGClient client = IGClient.builder().username("myuser").password("mylogin").login();
        
        InstagramDownloader dl = new InstagramDownloader();
        dl.downloadVideo("https://www.instagram.com/reel/CzeWZCYJ09R/", "C:\\temp");
    }
}

public class InstagramDownloader {

    private Document page;
    private final String USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/65.0.3325.181 Safari/537.36";

  public void downloadVideo(String url, String targetDirectory){
        String videoUrl = "";

        Helpers.validateURL(url);
        try {
            page = Jsoup.connect(url).userAgent(USER_AGENT).get();
            getAllElements(page);
            getAllNodes(page);
            //videoUrl = ???
            
        } catch (IOException e){
            e.printStackTrace();
        }      
       download(videoUrl, targetDirectory);
    }

 public void getAllElements(Document doc) {
         Elements children = new Elements();
         recurseOverElements(doc.getAllElements(), children);

         for (Element element : children) {
             System.out.println(element.tagName());
         }
    }

 public Elements recurseOverElements(Elements elementList, Elements children){
        if(elementList.size() == 0)
            return children;

        for (Element element : elementList) {

            recurseOverElements(element.children(), children);
            children.add(element);
        }
        return children;
    }
    
    public void getAllNodes(Document doc) {
        List<Node> allNodesInDom = new ArrayList<>();
        recurseOverNodes(doc.childNodes(), allNodesInDom);

        for (Node node : allNodesInDom) {
            System.out.println(node.nodeName());
        }
   }
    
    public List<Node> recurseOverNodes(List<Node> nodeList, List<Node> allChildNodeList){
        if(nodeList.size() == 0)
            return allChildNodeList;

        for (Node node : nodeList) {
            recurseOverNodes(node.childNodes(), allChildNodeList);
            allChildNodeList.add(node);
        }
        return allChildNodeList;
    }

private void download(String url, String targetDirectory){
        String[] tempName = url.split("/");
        String filename = tempName[tempName.length-1].split("[?]")[0];

        try(InputStream inputStream = URI.create(url).toURL().openStream()){
            int x = inputStream.read();
            System.out.println("x" + x);
            HttpURLConnection conn = (HttpURLConnection)URI.create(url).toURL().openConnection();
            Path targetPath = new File(targetDirectory + File.separator + filename).toPath();
            Files.copy(inputStream, targetPath, StandardCopyOption.REPLACE_EXISTING);

            int BYTES_PER_KB = 1024;
            double fileSize = ((double)conn.getContentLength() / BYTES_PER_KB);
        } catch (IOException e){
            e.printStackTrace();
        }
}

内容的提问来源于stack exchange,提问作者ecronin

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.06 08:10:18