使用ParserDelegator查找指定维基链接的流程控制问题求助
问题诊断与修复方案
核心问题分析
你遇到的「Checking children:」未打印的情况,大概率是以下两个原因之一:
callback.hasFound()逻辑错误:主页面解析后,hasFound()错误返回true,导致跳过子页面检查分支- 代码语法不完整:原代码的
try块、if分支及main方法均未闭合,编译或运行时会直接报错,根本执行不到子页面检查逻辑
分步修复方案
1. 补全代码基础语法
原代码存在语法缺失,先补全所有闭合括号,并添加异常处理避免程序崩溃:
public static void main(String args[]) throws Exception { String subject = args[0].replace(" ", "_"); System.out.println("Searching: " + subject + " - Wikipedia"); URL url = new URL("https://en.wikipedia.org/wiki/" + subject); try { BufferedReader in = new BufferedReader(new InputStreamReader(url.openStream())); StringBuilder sb = new StringBuilder(); String input; while ((input = in.readLine()) != null) { sb.append(input); } in.close(); ParserDelegator parser = new ParserDelegator(); MyParserCallback callback = new MyParserCallback(); parser.parse(new StringReader(sb.toString()), callback, true); // 增加调试输出,确认主页面是否找到目标 System.out.println("Main page found target: " + callback.hasFound()); if (!callback.hasFound()) { System.out.println("Checking children:"); Set<String> visitedLinks = new HashSet<>(); visitedLinks.add("/wiki/" + subject); for (String href : callback.getVisitedLinks()) { if (!visitedLinks.contains(href)) { visitedLinks.add(href); String childUrl = "https://en.wikipedia.org" + href; BufferedReader childIn = new BufferedReader(new InputStreamReader(new URL(childUrl).openStream())); StringBuilder childSb = new StringBuilder(); while ((input = childIn.readLine()) != null) { childSb.append(input); } childIn.close(); ParserDelegator childParser = new ParserDelegator(); MyParserCallback childCallback = new MyParserCallback(); childParser.parse(new StringReader(childSb.toString()), childCallback, true); if (childCallback.hasFound()) { System.out.println("Found target in child page: " + href); return; } } } } } catch (IOException e) { e.printStackTrace(); } }
2. 修正MyParserCallback逻辑
确保hasFound()能准确识别目标链接,且visitedLinks正确收集子页面链接:
class MyParserCallback extends DefaultHandler { private boolean foundTarget = false; private Set<String> wikiLinks = new HashSet<>(); @Override public void startElement(String uri, String localName, String qName, Attributes attributes) throws SAXException { if ("a".equals(qName)) { String href = attributes.getValue("href"); // 仅收集合法的维基百科子页面链接(排除非内容类链接,如Talk:、File:) if (href != null && href.startsWith("/wiki/") && !href.contains(":")) { wikiLinks.add(href); // 严格匹配目标链接 if ("/wiki/Geographic_coordinate_system".equals(href)) { foundTarget = true; } } } } public boolean hasFound() { return foundTarget; } public Set<String> getVisitedLinks() { return wikiLinks; } }
注意:原代码直接访问callback.visitedLinks的写法,若成员为私有会报错,需通过getVisitedLinks()方法获取。
3. 调试验证步骤
- 运行时传入
Linux作为参数,观察控制台输出的Main page found target:结果 - 如果主页面错误返回
true,检查MyParserCallback中目标链接的匹配逻辑(是否存在大小写、额外参数等问题) - 如果子页面链接为空,检查
wikiLinks的收集逻辑是否过滤了合法链接
内容的提问来源于stack exchange,提问作者Tiffany
相关产品推荐
相关产品推荐

