使用Jsoup解析NSE网页表格:无法获取tbody数据求助
解决Jsoup无法抓取NSE股票页面表格tbody内容的问题
问题背景
- 目标:提取NSE股票页面(ADANIENT)中
industryInfo表格的完整数据 - 问题:使用Jsoup直接请求页面,仅能获取thead部分,tbody内容为空
- 已尝试:查阅tbody解析相关问题,未找到有效解决方案
需要提取的表格结构
<table id="industryInfo" class="eq-series tbl-securityinfo cap-hide"> <caption></caption> <thead> <tr> <th>Macro-Economic Sector</th> <th>Sector</th> <th>Industry</th> <th>Basic Industry</th> </tr> </thead> <tbody class=""> <tr> <td>Commodities</td> <td>Metals & Mining</td> <td>Ferrous Metals</td> <td>Pig Iron</td> </tr> </tbody> </table>
当前使用的Jsoup代码
String url = "https://www.nseindia.com/get-quotes/equity?symbol=ADANIENT"; Document document = new Document(url); try { document = Jsoup.connect(url).userAgent("Mozilla/5.0").get(); } catch (IOException e) { e.printStackTrace(); } // System.out.println(document); Elements elements = document.select("#industryInfo"); for (Element element : elements) { System.out.println(element); }
问题原因
该页面的tbody内容并非静态HTML,而是通过前端JavaScript调用API动态渲染生成的。Jsoup仅能抓取初始返回的HTML源码,无法执行JS,因此获取不到动态加载的tbody数据。
解决方案
方案1:直接请求数据API
NSE的股票详情数据通过专用API接口返回,直接请求该接口获取JSON数据,解析效率更高。
import org.jsoup.Jsoup; import org.json.JSONObject; import java.io.IOException; public class NSEDataFetcher { public static void main(String[] args) { String apiUrl = "https://www.nseindia.com/api/quote-equity?symbol=ADANIENT"; try { String jsonResponse = Jsoup.connect(apiUrl) .userAgent("Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/118.0.0.0 Safari/537.36") .header("Referer", "https://www.nseindia.com/get-quotes/equity?symbol=ADANIENT") .ignoreContentType(true) .execute() .body(); JSONObject data = new JSONObject(jsonResponse); JSONObject industryInfo = data.getJSONObject("industryInfo"); System.out.println("Macro-Economic Sector: " + industryInfo.getString("macro")); System.out.println("Sector: " + industryInfo.getString("sector")); System.out.println("Industry: " + industryInfo.getString("industry")); System.out.println("Basic Industry: " + industryInfo.getString("basicIndustry")); } catch (IOException e) { e.printStackTrace(); } } }
注意:需引入JSON处理依赖(如
org.json),同时保持请求头与浏览器一致,避免被反爬拦截。
方案2:使用Selenium渲染JS页面
如果不想查找API,可通过Selenium启动真实浏览器,等待页面JS渲染完成后再抓取内容。
import org.jsoup.Jsoup; import org.jsoup.nodes.Document; import org.jsoup.nodes.Element; import org.jsoup.select.Elements; import org.openqa.selenium.WebDriver; import org.openqa.selenium.chrome.ChromeDriver; import java.time.Duration; public class SeleniumNSEScraper { public static void main(String[] args) { WebDriver driver = new ChromeDriver(); try { driver.get("https://www.nseindia.com/get-quotes/equity?symbol=ADANIENT"); driver.manage().timeouts().implicitlyWait(Duration.ofSeconds(10)); String pageSource = driver.getPageSource(); Document document = Jsoup.parse(pageSource); Elements rows = document.select("#industryInfo tbody tr"); for (Element row : rows) { Elements cells = row.select("td"); cells.forEach(cell -> System.out.print(cell.text() + "\t")); System.out.println(); } } finally { driver.quit(); } } }
注意:需配置对应版本的ChromeDriver,确保与本地Chrome浏览器版本匹配。
内容的提问来源于stack exchange,提问作者iCoder
相关产品推荐
相关产品推荐

