You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

JWikiDocs爬取维基百科仅下载种子页,无法批量爬取的问题求助

First, let's recap your setup to align on the context:

  • Running JWikiDocs on Ubuntu 17.10.1 (VirtualBox VM)
  • Compiled successfully via make clean + make all
  • Configured options.txt in data/Microsoft with totalPages=100 and seedURL=http://en.wikipedia.org/wiki/Microsoft
  • Executed crawler with: java -classpath lib/htmlparser/htmlparser.jar:lib/jwikidocs.jar jwikidocs.JWikiDocs -d data/Microsoft
  • Core issue: Only the seed page is downloaded; crawler terminates immediately without scraping internal links
  • Additional red flag: make test only updates the first document, and can't regenerate the full set of 100 test docs after deletion

Looking at your RetrievalLog.txt, the critical line is Current queue size: 0 right before processing the seed URL—this confirms the crawler isn't extracting any links from the seed page to add to its crawl queue. Here are actionable fixes to resolve this:

1. Validate the HTML Parser Dependency

JWikiDocs relies on htmlparser to extract links, so a broken or misconfigured dependency is a likely culprit:

  • Re-obtain the htmlparser.jar file (ensure it matches the version specified in JWikiDocs' documentation)
  • Fix potential path issues in your run command by using absolute paths instead of relative ones, to avoid classpath misalignment. For example:
    java -classpath /home/your-username/JWikiDocs/lib/htmlparser/htmlparser.jar:/home/your-username/JWikiDocs/lib/jwikidocs.jar jwikidocs.JWikiDocs -d /home/your-username/JWikiDocs/data/Microsoft
    

You adjusted maxDepth, but restrictive link filtering might be blocking all internal Wikipedia links:

  • Open Option.java and check for hardcoded filters that exclude valid Wikipedia article links (e.g., URLs starting with http://en.wikipedia.org/wiki/). Some crawlers default to excluding external links, anchors, or non-article pages—make sure these aren't overblocking.
  • Verify if there's an unconfigured allowedDomains setting in options.txt; many crawlers require explicit domain whitelisting to avoid wandering to external sites.

Add a quick debug step to confirm if links are being detected at all:

  • Modify the JWikiDocs source code to print extracted links to the console. For example, in the link-parsing section, add:
    System.out.println("Extracted link: " + link.getUrl());
    
  • Recompile with make all and run the crawler again. If no links appear in the terminal, the problem is definitely in the htmlparser integration or filtering logic.

4. Fix the make test Failure (Clue to Root Cause)

Since the test suite also fails to generate full results, this points to a core crawler logic issue, not just your configuration:

  • Delete the entire test directory, then re-run make test. Check the test's options.txt to confirm totalPages is set correctly and the seed URL is a page with abundant internal links.
  • If the test still only crawls one page, inspect the queue management code—links might be added but not persisted, or the queue is being cleared accidentally during processing.

5. Check for Wikipedia Anti-Crawling Measures

Even with increased sleep time, Wikipedia might be returning a modified page (like a captcha) that lacks links:

  • Manually fetch the seed URL from your VM to verify the content:
    curl http://en.wikipedia.org/wiki/Microsoft > test-page.html
    
  • Open test-page.html and confirm it contains normal article content with internal links. If it's a captcha or error page, add a valid User-Agent header to your crawler's requests (look for a setting in options.txt or source code to set a string like Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/117.0).

Since the crawler isn't populating its queue, focusing on link extraction and queue management will get you past this block. Let me know if any of these steps help narrow down the issue!

内容的提问来源于stack exchange,提问作者Scott

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.15 08:24:56