Java版Selenium多分隔符split()方法报错及基金规模爬取需求
问题说明
需要爬取页面的所有基金规模数据,去除₹、逗号、空格等分隔符后提取纯数字,求和后用TestNG断言做对比。当前代码运行时抛出java.lang.ArrayIndexOutOfBoundsException: Index 1 out of bounds for length 1错误。
爬取数据示例
爬取到的基金规模格式如下:
- ₹110 Crs
- ₹8,696 Crs
- ₹3,460 Crs
- ₹711 Crs
- ₹5,990 Crs
- ₹26,218 Crs
- ₹7,907 Crs
- ₹1,780 Crs
- ₹956 Crs
- ₹1,708 Crs
- ₹1,330 Crs
- ₹5,825 Crs
- ₹18,712 Crs
期望输出
110 8696 3460 711 5990 26218 7907 1780 956 1708 1330 5825 18712
现有错误代码
int total = 0; List<WebElement> totalFunds = driver.findElements(By.xpath("//div[@class='card-bordy']/div/div/div/div/div[2]/div/div/div[2]/div/p[2]")); for(int i=0;i<totalFunds.size();i++) { String funds = totalFunds.get(i).getText(); String fu = funds.split(" ")[0]; System.out.println(fu.split("₹")[1]); }
错误原因
数组越界的核心问题是:部分爬取到的文本中没有₹符号,或者通过split("₹")拆分后仅得到长度为1的数组,无法取到索引1的元素。同时现有代码未处理逗号分隔符,也缺少求和与断言逻辑。
解决方案
使用正则表达式一次性清理所有非数字字符,同时增加异常防护逻辑,实现数字提取、求和与TestNG断言:
import org.testng.Assert; int total = 0; List<WebElement> totalFunds = driver.findElements(By.xpath("//div[@class='card-bordy']/div/div/div/div/div[2]/div/div/div[2]/div/p[2]")); for (WebElement fundElement : totalFunds) { String fundsText = fundElement.getText().trim(); if (fundsText.isEmpty()) continue; // 跳过空文本元素 // 移除所有非数字字符,直接提取纯数字 String cleanAmount = fundsText.replaceAll("[^0-9]", ""); if (cleanAmount.isEmpty()) continue; // 防止清理后无有效数字 int amount = Integer.parseInt(cleanAmount); total += amount; System.out.println(amount); // 输出符合要求的纯数字格式 } // 示例数据总和为82403,替换为你的预期总和即可 int expectedTotal = 82403; Assert.assertEquals(total, expectedTotal, "基金规模总和不匹配");
代码说明
replaceAll("[^0-9]", ""):一次性移除₹、逗号、空格、Crs等所有非数字字符,比多次split更稳定- 增加空文本与空数字判断,避免解析崩溃
- 实现求和逻辑,并通过TestNG断言完成预期值对比
- 用增强for循环简化代码结构
内容的提问来源于stack exchange,提问作者Codeception
相关产品推荐
相关产品推荐

