Skip to content
Hosting Operations14 min read

DNS Not Resolving Not Working: Security Hardening Guide

Secure your DNS infrastructure with threat modeling, hardening steps, and audit checklists. Practical guide for hosting teams and website owners.

Written by Abdul AbrorTechnical Hosting Support Engineer
text
On this page

TL;DR — Key takeaways

  • DNS resolution failures often stem from security misconfigurations like restrictive firewalls blocking port 53, DNSSEC validation errors, or expired zone signatures that prevent legitimate queries from completing.
  • Hardening DNS requires layered defenses: rate limiting to prevent amplification attacks, DNSSEC for response authenticity, query logging for threat detection, and ACLs to restrict zone transfers and recursive queries.
  • Regular security audits should verify DNSSEC chain validity, check for open resolvers, review query logs for anomalies, and test failover behavior under simulated attack conditions.
  • Backup DNS infrastructure with geographically distributed nameservers and maintain offline DNSSEC key material in HSMs or encrypted storage to ensure recovery from compromise without extended downtime.

When DNS stops resolving, the root cause is often a security measure that has been misconfigured or has become too restrictive. Firewalls blocking legitimate queries, DNSSEC validation failures, and over-aggressive rate limiting can all manifest as resolution failures. For hosting teams and infrastructure engineers, understanding the security dimensions of DNS is essential to both preventing attacks and maintaining availability.

This guide covers the DNS threat model, provides a security audit checklist, walks through hardening steps for authoritative and recursive resolvers, and explains how to verify your configuration is secure without breaking legitimate traffic. Every recommendation includes rollback guidance and safe testing boundaries.

Understanding the DNS Threat Model

DNS infrastructure faces threats across multiple layers. Attackers exploit DNS for amplification attacks, cache poisoning, domain hijacking, data exfiltration, and reconnaissance. Each threat vector requires specific defenses.

Amplification attacks abuse open recursive resolvers to flood targets with large DNS responses. A small query can trigger a response 50-100 times larger, making DNS attractive for DDoS campaigns. Open resolvers also leak internal network information and can be used to bypass geographic restrictions.

Cache poisoning injects false records into resolver caches, redirecting users to malicious sites. DNSSEC provides cryptographic verification to prevent this, but implementation errors can cause resolution failures for legitimate domains.

Zone enumeration through AXFR transfers exposes your entire DNS database to attackers, revealing internal hostnames, IP ranges, and infrastructure topology. Unrestricted zone transfers are a critical misconfiguration.

DNS tunneling exfiltrates data or establishes command channels by encoding information in DNS queries and responses. This bypasses many firewall rules since DNS is rarely blocked outbound.

  • Amplification: Open resolvers used in reflection attacks
  • Poisoning: False records injected without DNSSEC protection
  • Enumeration: Zone transfers exposing infrastructure layout
  • Tunneling: Data exfiltration through encoded DNS queries
  • Hijacking: Registrar compromise or BGP attacks redirecting authoritative nameservers

DNS Security Audit Checklist

Run this audit quarterly and after any infrastructure changes. Document findings and track remediation status. Test during maintenance windows to avoid production impact.

Start by verifying your nameservers are not open resolvers. Query your server's IP from an external host with a domain you do not host. If it returns an answer, your resolver is open to the internet.

  • Test for open resolver: dig @your-ip example.com from external host; should refuse or timeout
  • Verify DNSSEC chain: dig +dnssec yourdomain.com; check for valid RRSIG records and AD flag
  • Check zone transfer restrictions: dig @your-ip yourdomain.com AXFR from unauthorized IP; should be refused
  • Review query rate limits: Send burst of queries and verify rate limiting activates without dropping legitimate traffic
  • Audit firewall rules: Confirm port 53 UDP/TCP open only to necessary sources; block AXFR from internet
  • Inspect query logs: Look for unusual patterns like high query rates from single IPs, queries for non-existent subdomains, or TXT record polling
  • Test DNSSEC validation: Query signed domain through your resolver; break signature and verify validation fails
  • Check software versions: Ensure DNS server software is current; known vulnerabilities exist in older BIND, Unbound, and PowerDNS releases
  • Verify secondary nameserver sync: Check SOA serial numbers match across all authoritative servers
  • Review access controls: Audit who can modify zone files and DNSSEC keys; require MFA for DNS management interfaces

Hardening Authoritative Nameservers

Authoritative servers answer queries for domains you host. Hardening focuses on preventing zone enumeration, ensuring availability, and validating responses with DNSSEC.

Restrict zone transfers to known secondary nameservers only. In BIND, use allow-transfer in named.conf with explicit IP addresses. Never use allow-transfer { any; }. Test by attempting AXFR from an unauthorized source and verify it is refused.

Enable DNSSEC signing for all zones. Generate zone-signing keys (ZSK) and key-signing keys (KSK) with appropriate algorithms. Use ECDSAP256SHA256 (algorithm 13) for new deployments. Automate key rotation with a 90-day ZSK lifetime and annual KSK rotation. Store KSK private keys offline or in hardware security modules.

Implement rate limiting to prevent your servers from being abused in amplification attacks. BIND's rate-limit statement and PowerDNS's max-qps-per-second setting limit responses to individual clients. Start conservatively at 10 queries per second per client and adjust based on legitimate traffic patterns.

Disable recursion entirely on authoritative servers. Set recursion no; in BIND. Authoritative servers should only answer for zones they host, never forward or resolve queries for other domains.

Hide server version information to reduce reconnaissance value. In BIND: version none; and hostname none;. This does not prevent attacks but removes easy fingerprinting.

    Hardening Recursive Resolvers

    Recursive resolvers query authoritative servers on behalf of clients. They require different hardening since they must perform recursion but should only serve authorized networks.

    Restrict recursion to trusted networks using ACLs. Define your internal networks in an ACL and apply it to allow-recursion. Block all external recursion attempts. Example BIND configuration: acl internals { 10.0.0.0/8; 192.168.0.0/16; }; options { allow-recursion { internals; }; };

    Enable DNSSEC validation to protect clients from cache poisoning. Set dnssec-validation auto; in BIND or val-permissive-mode no in Unbound. This causes the resolver to validate DNSSEC signatures and reject invalid responses. Monitor validation failures in logs as legitimate domains sometimes have broken DNSSEC.

    Implement response policy zones (RPZ) to block known malicious domains. RPZ acts as a DNS firewall, returning NXDOMAIN or a sinkhole address for blacklisted domains. Maintain RPZ feeds from threat intelligence sources and update daily.

    Configure query name minimization (RFC 7816) to reduce information leakage. This sends only necessary labels to each authoritative server instead of the full query. In Unbound: qname-minimisation yes;

    Set appropriate cache sizes and TTL limits. Oversized caches consume memory; undersized caches increase latency and upstream query load. Monitor cache hit rates and adjust. Limit maximum TTL to prevent poisoned records from persisting: max-cache-ttl 86400; in BIND limits to 24 hours.

    Enable logging for security monitoring but avoid logging every query due to performance and privacy concerns. Log refused queries, DNSSEC validation failures, and rate-limited clients. Rotate logs daily and retain for 30 days minimum for incident investigation.

      Firewall and Network-Level DNS Security

      DNS traffic must traverse firewalls correctly while preventing abuse. Misconfigurations at this layer commonly cause resolution failures that appear as DNS server problems.

      Allow UDP port 53 for standard queries but also allow TCP port 53 for large responses and zone transfers. DNSSEC responses often exceed 512 bytes and require TCP fallback. Blocking TCP 53 breaks DNSSEC validation.

      Implement stateful firewall rules that track DNS sessions and drop spoofed responses. Allow established connections and drop unexpected DNS packets. This prevents some cache poisoning attempts.

      Use firewall rate limiting as an additional defense layer. Limit inbound query rates from individual external sources to prevent contributing to amplification attacks even if server-level limits fail.

      For authoritative servers, allow queries from anywhere but restrict zone transfers (TCP port 53) to known secondary nameserver IPs only. For recursive resolvers, block all external access and allow queries only from internal networks.

      Consider DNS over HTTPS (DoH) or DNS over TLS (DoT) for client-to-resolver traffic if eavesdropping is a concern. This prevents network observation of queries but adds configuration complexity. DoH operates on TCP port 443; DoT on TCP port 853.

      Deploy geographically distributed nameservers with Anycast to improve availability and absorb DDoS attacks. Anycast routes queries to the nearest healthy server, automatically failing over during attacks or outages.

        Verifying DNS Security Configuration

        After applying hardening measures, verify they work correctly without breaking legitimate resolution. Test from multiple perspectives: internal clients, external resolvers, and unauthorized sources.

        Test DNSSEC validation with a tool like DNSViz or drill. Query your domain and verify the entire chain of trust from root to your zone. Look for valid RRSIG records, matching DS records at the parent, and the AD (authenticated data) flag in responses.

        Verify your servers are not open resolvers using online testing tools or by querying from an external host. Your server should refuse queries for domains you do not host when the source is not in your allowed recursion list.

        Simulate a zone transfer attempt from an unauthorized source: dig @your-authoritative-server yourdomain.com AXFR. This should be refused immediately. If it succeeds, your allow-transfer configuration is wrong.

        Test rate limiting by sending a burst of queries. Use dnsperf or similar load testing tools during a maintenance window. Verify that excessive queries are rate-limited without affecting normal traffic rates. Monitor logs to confirm rate-limit hits are logged.

        Check firewall rules with port scans and query attempts from outside your authorized networks. Verify TCP 53 is open for DNSSEC but zone transfers are blocked from untrusted sources.

        Review logs daily for the first week after hardening changes. Look for unexpected refusals, DNSSEC validation failures for legitimate domains, or legitimate clients hitting rate limits. Adjust thresholds based on actual traffic patterns.

          Common Hardening Pitfalls That Break Resolution

          Security measures can cause resolution failures if misconfigured. Understanding common mistakes helps you avoid them and troubleshoot faster when DNS stops working.

          Blocking TCP port 53 is the most common error. DNSSEC and large responses require TCP fallback. If you enable DNSSEC but block TCP 53, clients will see intermittent failures for signed domains. Solution: Always allow both UDP and TCP on port 53.

          Overly aggressive rate limiting drops legitimate queries during traffic spikes. Set rate limits based on actual traffic patterns, not theoretical minimums. Monitor rate-limit hit counts and adjust upward if legitimate clients are affected. Whitelist known high-volume sources like monitoring systems.

          DNSSEC validation failures occur when parent zones have stale DS records after a key rollback or when your resolver's trust anchor is outdated. Test DNSSEC thoroughly before enabling validation in production. Keep resolver software updated to maintain current root trust anchors.

          Restricting recursion to internal IPs breaks resolution if your ACL does not include all legitimate client networks. After infrastructure changes, verify your ACL covers new subnets. Use CIDR ranges rather than individual IPs to reduce maintenance.

          Expired DNSSEC signatures cause hard validation failures. Automate key rotation and signature refresh. Monitor signature expiration dates and alert at least 7 days before expiry. Unsigned zones are better than zones with expired signatures, which cause total resolution failure.

          Firewall rules that are too specific break during IP changes. For secondary nameservers, use hostnames in zone transfer ACLs when possible, though BIND requires IPs. Document nameserver IPs and update ACLs promptly when they change.

            Backup and Recovery Procedures

            DNS infrastructure must remain available during security incidents. Plan for compromise, misconfigurations, and attacks.

            Maintain offline backups of zone files, DNSSEC keys, and server configurations. Store encrypted backups in geographically separate locations. Test restoration quarterly to verify backups are valid and staff know the procedure.

            Keep DNSSEC key-signing key (KSK) private keys offline in a hardware security module (HSM) or encrypted USB drive in a safe. The ZSK handles routine signing and can be kept online with automated rotation. If the ZSK is compromised, use the offline KSK to sign a new ZSK without changing DS records at the parent.

            Document your DNS architecture and dependencies. Include nameserver IPs, upstream providers, registrar contacts, and parent zone contacts for DS record updates. During an incident, this documentation enables faster recovery.

            Run secondary nameservers with different providers or on different networks to ensure availability during attacks or provider outages. Use hidden primary configurations where the primary server is not publicly listed, reducing its attack surface.

            Plan for DNSSEC key compromise. Practice emergency key rollover procedures including generating new keys, signing zones, updating DS records, and monitoring propagation. The entire process takes 2-4 days due to TTLs, so act immediately when compromise is suspected.

            Establish out-of-band communication channels with your registrar and hosting provider. During a DNS hijacking attempt, you may not be able to use email under your domain. Have phone numbers, separate email accounts, and account credentials documented securely.

              Quick troubleshooting checklist

              • Test your nameservers for open resolver configuration from an external network
              • Verify DNSSEC chain validation for all hosted domains using DNSViz or drill
              • Confirm zone transfers are restricted to authorized secondary nameservers only
              • Enable query rate limiting on authoritative servers to prevent amplification abuse
              • Restrict recursive resolution to internal networks using ACLs
              • Enable DNSSEC validation on recursive resolvers to prevent cache poisoning
              • Allow both UDP and TCP on port 53 through firewalls for DNSSEC support
              • Implement response policy zones (RPZ) to block known malicious domains
              • Review DNS query logs for unusual patterns or potential tunneling activity
              • Automate DNSSEC key rotation with 90-day ZSK and annual KSK schedules
              • Store DNSSEC key-signing keys offline in encrypted storage or HSMs
              • Deploy geographically distributed secondary nameservers for high availability
              • Document nameserver IPs, registrar contacts, and parent zone contacts
              • Test zone transfer restrictions by attempting AXFR from unauthorized sources
              • Verify firewall rules block zone transfers from the internet while allowing queries
              • Monitor DNSSEC signature expiration dates and alert 7 days before expiry
              • Test backup restoration procedures quarterly to ensure recovery readiness
              • Review and update internal network ACLs after infrastructure changes
              • Enable query name minimization (RFC 7816) on recursive resolvers
              • Hide DNS server version information to reduce reconnaissance value
              • Maintain offline backups of zone files and configurations in separate locations

              FAQ

              Why does enabling DNSSEC cause DNS resolution to fail for some domains?

              DNSSEC validation failures occur when a domain's cryptographic signatures are invalid, expired, or the chain of trust is broken between the domain and the root. Common causes include expired zone signatures, stale DS records at the parent zone after a key rollback, outdated trust anchors in your resolver, or the parent zone not having DS records for a signed child zone. When DNSSEC validation is enabled, resolvers reject responses that fail validation, causing resolution to fail completely rather than returning potentially poisoned data. To fix this, verify the domain's DNSSEC chain using DNSViz, ensure your resolver's root trust anchor is current, and confirm both UDP and TCP port 53 are open since DNSSEC responses often require TCP fallback.

              How do I restrict DNS recursion without breaking legitimate client access?

              Restrict recursion by defining an ACL that lists your internal networks and applying it to the allow-recursion directive in your DNS server configuration. In BIND, create an ACL like 'acl internals { 10.0.0.0/8; 192.168.0.0/16; 172.16.0.0/12; };' covering your private IP ranges, then set 'allow-recursion { internals; };' in the options block. This permits recursion only from listed networks while refusing external queries. For authoritative servers, disable recursion entirely with 'recursion no;'. After configuration changes, test from both an internal client and an external host to verify internal clients can resolve any domain while external queries are refused. Update ACLs promptly when adding new internal subnets to avoid breaking access for new infrastructure.

              What is the safest way to implement DNS rate limiting without dropping legitimate traffic?

              Start with conservative rate limits based on your actual traffic patterns rather than theoretical minimums. Monitor DNS query rates per client IP for one week to establish a baseline, then set initial limits at 2-3 times your observed peak rates. In BIND, use 'rate-limit { responses-per-second 10; window 5; };' as a starting point, which allows 10 responses per second per client averaged over 5 seconds. Enable detailed logging of rate-limited queries and review daily for the first two weeks. If legitimate clients like monitoring systems or busy web servers hit limits, whitelist their IPs using 'exempt-clients' or increase the global limit. Adjust thresholds incrementally based on observed behavior rather than making large changes. For public authoritative servers, focus rate limiting on error responses (NXDOMAIN, SERVFAIL) which are more commonly abused in amplification attacks while being more permissive with valid responses.