如何配置WordPress:允许Fess爬取,拒绝Google爬取
Got it, let's break this down. When you check that "Discourage Search Engine from indexing this site" option in WordPress, it does two key things: it adds a Disallow: / rule to your robots.txt file, and injects a <meta name="robots" content="noindex, nofollow"> tag into all your pages. To let Fess crawl your site while keeping other search engines out, you have two solid approaches:
Option 1: Configure Fess to Ignore Robots.txt Restrictions
This is the quickest fix if you don't want to tweak WordPress code. Fess follows robots.txt rules by default, so you can disable that for your target crawler:
- Log into your Fess admin dashboard, go to Crawler > Crawler Settings.
- Either edit your existing crawler configuration or create a new one for your WordPress site.
- Scroll to the Advanced Settings section, find the "Follow Robots.txt" option, and set it to No.
- Save the changes, then restart the crawler. Fess will now ignore the
Disallowrule in your robots.txt and crawl your site.
⚠️ Note: If you have security plugins like Wordfence installed on WordPress, make sure to add Fess's user agent (Fess) or its server IP to your allowed list to avoid being blocked.
Option 2: Make WordPress Exempt Fess from Noindex Rules
This is a more precise approach—it keeps the noindex/nofollow rules for all other crawlers, but lets Fess through. You can do this with code or a plugin:
Code Method (Add to Your Theme's functions.php)
Edit your active theme's functions.php file (or use a child theme to avoid losing changes on updates) and add this snippet:
add_filter('wp_robots', 'allow_fess_crawling'); function allow_fess_crawling($robots) { $user_agent = $_SERVER['HTTP_USER_AGENT']; // Check if the crawler is Fess if (strpos($user_agent, 'Fess') !== false) { // Remove noindex and nofollow directives for Fess unset($robots['noindex']); unset($robots['nofollow']); } return $robots; }
This code hooks into WordPress's robot meta tag generation and removes the noindex/nofollow rules only when the crawler's user agent is Fess.
Plugin Method (Using Yoast SEO)
If you prefer not to edit code, use Yoast SEO (a popular WordPress SEO plugin):
- Install and activate Yoast SEO if you haven't already.
- Go to Yoast SEO > Settings > Advanced > Robots.txt.
- In the "Allow search engines to crawl specific URLs" section, add a rule that allows Fess's user agent. For example:
User-agent: Fess Allow: / - Save the settings. Yoast will override the default robots.txt rules for Fess, letting it crawl your entire site.
Either method should work—pick the one that fits your comfort level with code or plugins.
内容的提问来源于stack exchange,提问作者久保圭司

