使用Jsoup无法获取Instagram的video标签,求解决方法
Instagram视频下载问题:无法通过Jsoup定位视频资源
我是后端开发者,已有15年未接触HTML文档解析,对相关操作不太熟悉;同时我对Instagram的运作机制也处于学习阶段。
我尝试下载Instagram上的视频,按常规逻辑视频应该在<video>标签中,但无论通过何种方式遍历org.jsoup.nodes.Document的子元素,都无法识别到该标签。我试过调用Document.children().select(*)方法,怀疑Instagram隐藏了视频源,但毫无头绪。
我原本预期页面会存在og:video元标签(页面中title、img等元标签均正常存在),并尝试通过代码page.select("meta[property=og:video]").first().attr("content");获取视频地址,但该元标签并不存在。我还在InstagramDownloader类中编写了两个递归方法遍历所有节点和元素(方法来自另一Stack Overflow问题),但依然没找到获取视频的线索。我甚至不确定即便拿到视频的src URL,是否能成功完成下载。
相关代码如下:
public class Application { public static void main(String[] args) { try { login(); } catch (Exception e) { e.printStackTrace(); } } public static void login() throws IGLoginException, InterruptedException, ExecutionException{ IGClient client = IGClient.builder().username("myuser").password("mylogin").login(); InstagramDownloader dl = new InstagramDownloader(); dl.downloadVideo("https://www.instagram.com/reel/CzeWZCYJ09R/", "C:\\temp"); } } public class InstagramDownloader { private Document page; private final String USER_AGENT = "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/65.0.3325.181 Safari/537.36"; public void downloadVideo(String url, String targetDirectory){ String videoUrl = ""; Helpers.validateURL(url); try { page = Jsoup.connect(url).userAgent(USER_AGENT).get(); getAllElements(page); getAllNodes(page); //videoUrl = ??? } catch (IOException e){ e.printStackTrace(); } download(videoUrl, targetDirectory); } public void getAllElements(Document doc) { Elements children = new Elements(); recurseOverElements(doc.getAllElements(), children); for (Element element : children) { System.out.println(element.tagName()); } } public Elements recurseOverElements(Elements elementList, Elements children){ if(elementList.size() == 0) return children; for (Element element : elementList) { recurseOverElements(element.children(), children); children.add(element); } return children; } public void getAllNodes(Document doc) { List<Node> allNodesInDom = new ArrayList<>(); recurseOverNodes(doc.childNodes(), allNodesInDom); for (Node node : allNodesInDom) { System.out.println(node.nodeName()); } } public List<Node> recurseOverNodes(List<Node> nodeList, List<Node> allChildNodeList){ if(nodeList.size() == 0) return allChildNodeList; for (Node node : nodeList) { recurseOverNodes(node.childNodes(), allChildNodeList); allChildNodeList.add(node); } return allChildNodeList; } private void download(String url, String targetDirectory){ String[] tempName = url.split("/"); String filename = tempName[tempName.length-1].split("[?]")[0]; try(InputStream inputStream = URI.create(url).toURL().openStream()){ int x = inputStream.read(); System.out.println("x" + x); HttpURLConnection conn = (HttpURLConnection)URI.create(url).toURL().openConnection(); Path targetPath = new File(targetDirectory + File.separator + filename).toPath(); Files.copy(inputStream, targetPath, StandardCopyOption.REPLACE_EXISTING); int BYTES_PER_KB = 1024; double fileSize = ((double)conn.getContentLength() / BYTES_PER_KB); } catch (IOException e){ e.printStackTrace(); } }
内容的提问来源于stack exchange,提问作者ecronin
相关产品推荐
相关产品推荐

