ShipworkSite inspection
Checks

Crawlability

PaginationAre paginated pages set up right?Internal searchIs site search eating crawl budget?URL parametersAre parameters creating duplicates?Log analyserWhat does Googlebot actually crawl?Robots builderNeed a robots.txt file?Soft 404sAre pages “not found” but returning 200?JavaScriptCan crawlers see it without running JS?FreshnessIs my site quietly going stale?Robots.txtDoes robots.txt say what I think?

Indexing

Sitemap lastmodAre my sitemap dates valid?SitemapIs my sitemap actually fetchable?IndexingIs Google allowed to index my pages?CanonicalsIs Google indexing the wrong URL?Index signalsDo sitemap and index signals agree?Sitemap generatorNeed a sitemap.xml for my site?Bulk URL statusWhere does each of these URLs really land?Redirect builderNeed the redirect rules?RedirectsAre my URLs answering directly, over HTTPS?

On-page

Keyword ideasWhat are people searching for?SERP previewHow does my page look in Google?Content qualityAre your pages too thin or too alike?Image weightAre images slowing the page?CannibalizationAre pages competing with each other?On-page checkDoes the page use its target phrase?ImagesAre my images accessible and loading?Meta tag builderWhat should my title and share tags say?DuplicatesDo my pages compete for one query?

Links

Orphan pagesWhich pages can no link reach?Link graphHow deep do your pages sit?Anchor textDo links say what they point at?Outbound linksAre external links still alive?Broken linksAre internal links sending visitors nowhere?

Structured data

Rich resultsIs my markup eligible for a rich result?Structured dataIs my product schema valid?Schema coverageDo my key pages carry structured data?Social previewHow does my page look when shared?Schema vs pageDoes markup match the page price?Schema builderNeed valid JSON-LD?

International

Hreflang sitemapDo page and sitemap hreflang agree?HreflangDo my language versions link back?Hreflang builderNeed the hreflang tags?Page DoctorWhy is this page not doing well?Redirect planMy site has dead links — what redirects do I write?Sitemap diffIs anything missing from my sitemap?Robots simulatorWhat does my robots.txt actually block?Can Googlebot?Can Googlebot fetch this URL?
Every error, explainedGuides
Pricing Learn Error guides Run free audit

+1 free check

Welcome to Shipwork

Sign in free and your next check this month is on us.

Continue with GoogleContinue with email

We email you a 6-digit code. No password.

No card needed. Reports and watches follow you to any device.

X-Robots-Tag noindex: find it, change it and verify the response

An X-Robots-Tag: noindex header tells crawlers not to include a fetched URL in search results. Find it in the HTTP response, identify which Nginx or upstream rule adds it, and remove only the rule that is no longer intended. Then let crawlers fetch the URL and verify the uncached public response. A robots.txt block can prevent them from seeing the header at all.

Published

Check the actual HTTP response

Use a response-header request to look for X-Robots-Tag. A HEAD request is a quick first check:

curl -sSI https://example.test/reports/monthly.pdf

Look through the response headers for X-Robots-Tag: noindex, including any crawler-specific value such as googlebot: noindex. If the server, proxy or application handles HEAD differently from GET, repeat with a GET while discarding the body:

curl -sS -D - -o /dev/null https://example.test/reports/monthly.pdf

Use the exact public URL, including its scheme, hostname and path. If it redirects, inspect each response and then check the final URL too. A header on the first redirect and a header on the destination are separate observations. Also compare a URL you expect to be indexable with one you intentionally exclude; that can reveal a broad rule affecting more paths than planned.

Google’s robots meta tag and X-Robots-Tag documentation explains that the HTTP header can express the same robots directives as an HTML robots meta tag, and can apply to non-HTML resources such as PDFs. The directive is visible only when a crawler can fetch the response. It does not provide access control for private documents.

If the response has no X-Robots-Tag, inspect the HTML for a <meta name="robots" content="noindex"> tag and check for crawler-specific tags too. A clean HTTP header does not rule out a meta directive, and a meta tag check does not rule out an HTTP header.

Add an X-Robots-Tag rule in Nginx

To noindex one exact PDF while leaving other files alone, a location like this can add the response header:

location = /reports/monthly.pdf {
    add_header X-Robots-Tag "noindex";
}

For a file family, a case-insensitive regular-expression location can match the intended extensions:

location ~* \.(pdf|docx)$ {
    add_header X-Robots-Tag "noindex";
}

Place the rule in the correct server context and consider how existing locations are selected. This is an example, not a drop-in replacement for every site’s config: a more specific location, an upstream application or another proxy may determine the final response. Nginx’s add_header normally applies to a defined set of response codes; adding always changes that behavior, so choose it only if you need the header on other response statuses too.

There is an inheritance detail worth checking before placing the directive in a nested location. Under Nginx’s default rules, add_header directives from a parent level are inherited only when the child level has no add_header directives of its own. A new location-level header can therefore affect other headers you expected to inherit. Review the active config and version-specific behavior before using newer inheritance controls. The Nginx headers module documentation describes the directive’s context, response-code behavior and inheritance.

Remove the right rule when noindex is unintended

First find the source. Search the active configuration and included files for X-Robots-Tag, then check the application, hosting panel, CDN rules and generated configuration. A response may contain multiple values from different layers. Removing one Nginx directive will not remove a header that an upstream app or edge proxy still sends.

If only a temporary maintenance or preview state added the rule, remove that rule from the source that owns it. Do not simply add a second header with a less restrictive value and assume it cancels the first. Google documents that when robots directives conflict, it follows the more restrictive rule; a remaining noindex can keep the page excluded.

Before changing the config, preserve the current file and identify a nearby URL that should retain its existing policy. After the edit, validate the complete configuration with nginx -t, apply it through your normal process and inspect both the formerly noindexed URL and the control URL. If the rule is generated by a deployment template, change the template or setting that generates it; a manual edit may disappear on the next release.

Also check the HTML robots meta tag, crawler-specific meta tags, canonical URL and response status. If the site intentionally excludes a page, keep the intended rule in place. Removing noindex from the wrong page merely trades an exclusion for a possible duplicate, private-content or low-value indexing problem.

Verify crawl access and cached responses

After removal, check the response again from the public URL with GET and, if useful, HEAD. Confirm that no response in the redirect chain or final document still carries a noindex header. Fetch the HTML and inspect its robots meta tags as well. Compare the response at the origin and edge if the site has a cache or proxy, and purge or wait out a stale cached response according to that layer’s normal process.

Keep a short before-and-after record for the exact URL: request method, status, redirect destination, X-Robots-Tag value, meta robots value and the time checked. If a CDN exposes cache status or age headers, save those too. That evidence helps separate a rule that remains active from a response that was cached before the fix, and gives the next deploy reviewer a known-good URL to recheck.

Make sure the URL is not disallowed in robots.txt. Google says it must crawl a URL to discover the noindex directive; when crawling is blocked, the crawler cannot read the response header or page meta tag. A robots block and a noindex tag answer different questions, so do not add a robots disallow as a substitute for excluding a URL from search. See Google’s robots.txt guide and the existing robots.txt vs. noindex explanation.

Use the indexing check to inspect the public HTML, response headers and robots rules for a set of key pages. The bulk URL status checker also reports noindex from the final HTML or X-Robots-Tag header when it receives a successful HTML response. These are checks of what the request observed, not proof that Google has already refreshed its index. After the technical signals are corrected, use Search Console URL Inspection or indexing reports to review Google’s own processing state.

If you removed a deliberate temporary exclusion, monitor the response and page source after the next cache refresh and deployment. Confirm the page returns the expected status, the intended canonical, and no leftover robots directive. Search visibility may take time to reflect the updated response; a single clean request does not force immediate re-indexing.

Find out if this is happening to you

Shipwork reads public page HTML, response headers and robots rules to identify noindex and crawl signals; it cannot force search systems to refresh an index. Free, no account, no signup. Paste your store address.

Check indexing signals

Questions

Can I check X-Robots-Tag with HEAD?
Yes, HEAD can show response headers without returning the body. If the application or proxy treats HEAD differently, repeat with GET and discard the body.
Does robots.txt remove a noindex header?
No. A robots.txt block can prevent a crawler from fetching the response that contains the noindex rule, so it may not discover the directive.
Why is noindex still present after I removed it in Nginx?
Another layer may be adding it, such as an upstream app, CDN, included config or cached response. Inspect the public response and trace each header source.
Does removing noindex make Google index the page immediately?
No. It removes one indexing instruction from the response. Crawling, processing and indexing are determined by the search system and may take time.

Keep reading