Check the actual HTTP response
Use a response-header request to look for X-Robots-Tag. A HEAD request is a quick first check:
curl -sSI https://example.test/reports/monthly.pdf
Look through the response headers for X-Robots-Tag: noindex, including any crawler-specific value such as googlebot: noindex. If the server, proxy or application handles HEAD differently from GET, repeat with a GET while discarding the body:
curl -sS -D - -o /dev/null https://example.test/reports/monthly.pdf
Use the exact public URL, including its scheme, hostname and path. If it redirects, inspect each response and then check the final URL too. A header on the first redirect and a header on the destination are separate observations. Also compare a URL you expect to be indexable with one you intentionally exclude; that can reveal a broad rule affecting more paths than planned.
Google’s robots meta tag and X-Robots-Tag documentation explains that the HTTP header can express the same robots directives as an HTML robots meta tag, and can apply to non-HTML resources such as PDFs. The directive is visible only when a crawler can fetch the response. It does not provide access control for private documents.
If the response has no X-Robots-Tag, inspect the HTML for a <meta name="robots" content="noindex"> tag and check for crawler-specific tags too. A clean HTTP header does not rule out a meta directive, and a meta tag check does not rule out an HTTP header.
Add an X-Robots-Tag rule in Nginx
To noindex one exact PDF while leaving other files alone, a location like this can add the response header:
location = /reports/monthly.pdf {
add_header X-Robots-Tag "noindex";
}For a file family, a case-insensitive regular-expression location can match the intended extensions:
location ~* \.(pdf|docx)$ {
add_header X-Robots-Tag "noindex";
}Place the rule in the correct server context and consider how existing locations are selected. This is an example, not a drop-in replacement for every site’s config: a more specific location, an upstream application or another proxy may determine the final response. Nginx’s add_header normally applies to a defined set of response codes; adding always changes that behavior, so choose it only if you need the header on other response statuses too.
There is an inheritance detail worth checking before placing the directive in a nested location. Under Nginx’s default rules, add_header directives from a parent level are inherited only when the child level has no add_header directives of its own. A new location-level header can therefore affect other headers you expected to inherit. Review the active config and version-specific behavior before using newer inheritance controls. The Nginx headers module documentation describes the directive’s context, response-code behavior and inheritance.
Remove the right rule when noindex is unintended
First find the source. Search the active configuration and included files for X-Robots-Tag, then check the application, hosting panel, CDN rules and generated configuration. A response may contain multiple values from different layers. Removing one Nginx directive will not remove a header that an upstream app or edge proxy still sends.
If only a temporary maintenance or preview state added the rule, remove that rule from the source that owns it. Do not simply add a second header with a less restrictive value and assume it cancels the first. Google documents that when robots directives conflict, it follows the more restrictive rule; a remaining noindex can keep the page excluded.
Before changing the config, preserve the current file and identify a nearby URL that should retain its existing policy. After the edit, validate the complete configuration with nginx -t, apply it through your normal process and inspect both the formerly noindexed URL and the control URL. If the rule is generated by a deployment template, change the template or setting that generates it; a manual edit may disappear on the next release.
Also check the HTML robots meta tag, crawler-specific meta tags, canonical URL and response status. If the site intentionally excludes a page, keep the intended rule in place. Removing noindex from the wrong page merely trades an exclusion for a possible duplicate, private-content or low-value indexing problem.
Verify crawl access and cached responses
After removal, check the response again from the public URL with GET and, if useful, HEAD. Confirm that no response in the redirect chain or final document still carries a noindex header. Fetch the HTML and inspect its robots meta tags as well. Compare the response at the origin and edge if the site has a cache or proxy, and purge or wait out a stale cached response according to that layer’s normal process.
Keep a short before-and-after record for the exact URL: request method, status, redirect destination, X-Robots-Tag value, meta robots value and the time checked. If a CDN exposes cache status or age headers, save those too. That evidence helps separate a rule that remains active from a response that was cached before the fix, and gives the next deploy reviewer a known-good URL to recheck.
Make sure the URL is not disallowed in robots.txt. Google says it must crawl a URL to discover the noindex directive; when crawling is blocked, the crawler cannot read the response header or page meta tag. A robots block and a noindex tag answer different questions, so do not add a robots disallow as a substitute for excluding a URL from search. See Google’s robots.txt guide and the existing robots.txt vs. noindex explanation.
Use the indexing check to inspect the public HTML, response headers and robots rules for a set of key pages. The bulk URL status checker also reports noindex from the final HTML or X-Robots-Tag header when it receives a successful HTML response. These are checks of what the request observed, not proof that Google has already refreshed its index. After the technical signals are corrected, use Search Console URL Inspection or indexing reports to review Google’s own processing state.
If you removed a deliberate temporary exclusion, monitor the response and page source after the next cache refresh and deployment. Confirm the page returns the expected status, the intended canonical, and no leftover robots directive. Search visibility may take time to reflect the updated response; a single clean request does not force immediate re-indexing.
Shipwork reads public page HTML, response headers and robots rules to identify noindex and crawl signals; it cannot force search systems to refresh an index. Free, no account, no signup. Paste your store address.
Check indexing signalsQuestions
Can I check X-Robots-Tag with HEAD?
Does robots.txt remove a noindex header?
Why is noindex still present after I removed it in Nginx?
Does removing noindex make Google index the page immediately?
Keep reading