JWikiDocs爬取维基百科仅下载种子页,无法批量爬取的问题求助
First, let's recap your setup to align on the context:
- Running JWikiDocs on Ubuntu 17.10.1 (VirtualBox VM)
- Compiled successfully via
make clean+make all - Configured
options.txtindata/MicrosoftwithtotalPages=100andseedURL=http://en.wikipedia.org/wiki/Microsoft - Executed crawler with:
java -classpath lib/htmlparser/htmlparser.jar:lib/jwikidocs.jar jwikidocs.JWikiDocs -d data/Microsoft - Core issue: Only the seed page is downloaded; crawler terminates immediately without scraping internal links
- Additional red flag:
make testonly updates the first document, and can't regenerate the full set of 100 test docs after deletion
Looking at your RetrievalLog.txt, the critical line is Current queue size: 0 right before processing the seed URL—this confirms the crawler isn't extracting any links from the seed page to add to its crawl queue. Here are actionable fixes to resolve this:
1. Validate the HTML Parser Dependency
JWikiDocs relies on htmlparser to extract links, so a broken or misconfigured dependency is a likely culprit:
- Re-obtain the htmlparser.jar file (ensure it matches the version specified in JWikiDocs' documentation)
- Fix potential path issues in your run command by using absolute paths instead of relative ones, to avoid classpath misalignment. For example:
java -classpath /home/your-username/JWikiDocs/lib/htmlparser/htmlparser.jar:/home/your-username/JWikiDocs/lib/jwikidocs.jar jwikidocs.JWikiDocs -d /home/your-username/JWikiDocs/data/Microsoft
2. Audit Link Filtering Logic
You adjusted maxDepth, but restrictive link filtering might be blocking all internal Wikipedia links:
- Open
Option.javaand check for hardcoded filters that exclude valid Wikipedia article links (e.g., URLs starting withhttp://en.wikipedia.org/wiki/). Some crawlers default to excluding external links, anchors, or non-article pages—make sure these aren't overblocking. - Verify if there's an unconfigured
allowedDomainssetting inoptions.txt; many crawlers require explicit domain whitelisting to avoid wandering to external sites.
3. Debug Link Extraction
Add a quick debug step to confirm if links are being detected at all:
- Modify the JWikiDocs source code to print extracted links to the console. For example, in the link-parsing section, add:
System.out.println("Extracted link: " + link.getUrl()); - Recompile with
make alland run the crawler again. If no links appear in the terminal, the problem is definitely in the htmlparser integration or filtering logic.
4. Fix the make test Failure (Clue to Root Cause)
Since the test suite also fails to generate full results, this points to a core crawler logic issue, not just your configuration:
- Delete the entire test directory, then re-run
make test. Check the test'soptions.txtto confirmtotalPagesis set correctly and the seed URL is a page with abundant internal links. - If the test still only crawls one page, inspect the queue management code—links might be added but not persisted, or the queue is being cleared accidentally during processing.
5. Check for Wikipedia Anti-Crawling Measures
Even with increased sleep time, Wikipedia might be returning a modified page (like a captcha) that lacks links:
- Manually fetch the seed URL from your VM to verify the content:
curl http://en.wikipedia.org/wiki/Microsoft > test-page.html - Open
test-page.htmland confirm it contains normal article content with internal links. If it's a captcha or error page, add a valid User-Agent header to your crawler's requests (look for a setting inoptions.txtor source code to set a string likeMozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:109.0) Gecko/20100101 Firefox/117.0).
Since the crawler isn't populating its queue, focusing on link extraction and queue management will get you past this block. Let me know if any of these steps help narrow down the issue!
内容的提问来源于stack exchange,提问作者Scott

