新DNS服务器内部解析故障求助:间歇性失效需Ping触发恢复
Hey Dave, this is a tricky but super common issue when migrating DNS servers—let’s walk through the most likely culprits and fixes based on what you’ve described:
1. Check if Your Internal Zone is Properly Loaded & Persisted
First, let’s rule out the basics: when your internal resolution dies, does your DNS server actually have the internal zone data in its memory?
- On most DNS servers (like BIND), run
rndc statusto see if the internal zone is listed as "loaded". If it’s missing, that means the server isn’t keeping the zone loaded permanently. - Compare your old server’s zone config to the new one. Did you miss a parameter like
allow-transfer(if it’s a slave) orauto-dnssec maintain? Old servers sometimes have legacy settings that keep zones cached longer, while new defaults might unload inactive zones to save memory.
2. Verify Zone File Permissions & Access
It’s easy to overlook permissions when copying configs:
- Make sure the DNS service user (e.g.,
namedon Linux) has read access to your internal zone file. If it can’t read the file during startup, it might only load it when forced (like when you ping, which triggers a local query that forces the server to recheck the zone). - Run
named-checkzone your.internal.domain /path/to/zone/fileto validate the zone file syntax—even a tiny typo that the old server ignored (due to looser version defaults) could break loading on the new server.
3. Dig Into DNS Server Logs for Clues
Logs will tell you exactly what’s going wrong when resolution fails:
- Tail your DNS service’s log file (e.g.,
tail -f /var/log/named/named.logfor BIND) and wait for the resolution to break. Look for errors like:zone your.internal.domain/IN: loading from master file failed: permission deniedzone your.internal.domain/IN: not loaded due to errorscache cleanup removed zone data for your.internal.domain
These errors will point directly to whether it’s a permission, syntax, or cache expiration issue.
4. Test Manual Zone Reload vs. Ping Trigger
Instead of pinging, try manually reloading the zone with rndc reload your.internal.domain when resolution fails. If this fixes it immediately, that confirms the server isn’t automatically refreshing the zone.
- Check your zone’s
refreshandexpiretimers in the SOA record. If the refresh interval is too long, the server might wait too long to recheck the zone, and when cache expires, it drops the data. Old servers might have had a shorter refresh time that you didn’t copy over.
5. Rule Out Cache Eviction Issues
If your suspicion about cache is correct, check if the new server has a lower cache limit than the old one:
- For BIND, look at the
max-cache-sizeparameter innamed.conf. If it’s set too low, the server might evict the less-frequently accessed internal zone data to make space for public DNS records. Increasing this limit (or setting it tounlimitedtemporarily for testing) could fix the problem.
Quick Validation Test
Next time resolution fails, run dig @localhost your.internal.host directly on the DNS server. If this returns no results, but dig your.internal.host (using the server’s own resolver) works afterward, that means the server isn’t serving the zone data until it’s forced to query it itself. This confirms a zone loading or persistence issue, not just client-side cache.
内容的提问来源于stack exchange,提问作者Dave Lamb

