Skip to content
holala.ai is live!AI image generation ↗
Prix Studio

PRIX STUDIO / JOURNAL

What Is robots.txt? Rules, Creation and Validation Guide

Robots.txt tells cooperating crawlers which URL paths they may request. Before writing it, identify the problem: unnecessary crawling, visibility in search results and unauthorized access require different controls.

Prix Studio7 min readUpdated
What Is robots.txt? Rules, Creation and Validation Guide
Prix Studio · AI-assisted editorial illustration
01

Crawling, indexing and privacy require different decisions

A technical SEO audit considers robots.txt as part of crawl accessibility. It controls requests from cooperating crawlers such as Googlebot. It does not directly improve rankings or physically stop every bot. A small, straightforward website may need no custom restrictions; for large URL collections, inspect server logs before deciding which requests are unnecessary.

A blocked HTML or PDF URL can still be discovered through links and appear in search results. If Google cannot fetch the page, it cannot read its noindex instruction. Do not interpret a crawl restriction as confirmation that a URL has disappeared from the index.

Reduce unnecessary requests

Consider Disallow for defined URL paths. Check that restrictions preserve product discovery and resources needed to understand pages.

Prevent search visibility

Use meta robots noindex on an accessible page or an appropriate X-Robots-Tag response header for a PDF. The crawler must be able to read the instruction.

Keep content private

Enforce authentication and access permissions. Robots.txt is public; do not include credentials or private document locations. RFC 9309 explicitly separates it from content security.

02

Read rules against representative URLs

When applying a technical SEO checklist, record one URL expected to pass and another expected to be blocked for each rule. User-agent identifies the crawler group, Disallow restricts paths and Allow defines an exception. Field names are case-insensitive; path values are case-sensitive. Comments start with # rather than an explanatory sentence in parentheses after a command.

Google selects the most specific matching crawler group without combining it with the general * group. Multiple groups for that specific crawler are combined. The longest matching path rule determines the result; equally conflicting rules use the less restrictive instruction. The last line does not automatically win.

ExampleMeaning and check
User-agent: *General group for crawlers without their own matching group.
Disallow: /admin/Restricts this path prefix; it does not enforce login permissions.
Allow: /admin/public.cssMore specific exception; confirm that the resource is required.
Disallow: /Blocks all paths for the group; intentional use needs a clear purpose.
Disallow:An empty value defines no restricted path.
Disallow: /*.pdf$Matches paths ending in .pdf; a version with ?id=1 is different.
* and $* matches a character sequence and $ marks the end. Avoid indiscriminately blocking every query string.

FROM READING TO A NEXT STEP

Review rules against your URL examples

Share the current robots.txt, sitemaps, important product or service paths and crawl issues so we can define the change and its validation steps.

Discuss technical SEO scope ↗
03

Check the address and response as well as the file

Your hosting arrangements determine which layer serves the file. The expected location is at the root, such as https://example.com/robots.txt; /assets/robots.txt is not equivalent. HTTPS, HTTP, www, another subdomain and a nonstandard port can have separate scopes. Sharing a server does not make their rules identical.

Prepare UTF-8 plain text. A CMS can generate the response dynamically, so a physical public_html file or FTP access is not always required. Check for an unexpected login screen, HTML error response, stale CDN cache or redirect to the wrong host. Google processes the first 500 KiB; consolidate unnecessary individual rules.

Google treats 4xx responses other than 429 as an absent robots.txt file. Returning 403 for robots.txt is not a way to prohibit all crawling. Server and network errors have different caching and retry behavior. A failed fetch is not equivalent to a valid Disallow rule. Google’s specification guide explains these response distinctions.

04

Assign one owner for platform changes

Your CMS choice affects robots.txt management. WordPress can generate its default response dynamically; the official do_robots source restricts the administration path and allows admin-ajax.php. Identify whether a physical file, plugin or server rule controls the published response. Using Yoast SEO or Rank Math does not automatically make the result correct.

Shopify provides default rules suitable for most stores. When customization is necessary, use robots.txt.liquid while preserving the platform’s evolving default groups. Replacing the entire output with static text can miss later updates. Follow the current Shopify guidance rather than an old admin screenshot.

WooCommerce cart, checkout and account paths depend on language and permalink settings. An English template may not match a Turkish store. For ikas and IdeaSoft, verify the current panel, published file and support scope instead of assuming identical editing permissions across plans. Blocking /en/ or /de/ content paths indiscriminately can affect discovery of translated pages.

Cotexlab, a selected Prix Studio website
Cotexlab · A reference from our website portfolio Selected work ↗
05

Separate filters, pagination and crawler purposes

As with deciding whether to update or remove old content, assign a purpose to each URL group. Evaluate crawling and indexing separately for cart, thank-you, internal search and account pages. A PDF may contain a useful manual or specification. Review the value of tag and author archives before blocking them collectively.

Inventory representative filter and sorting URLs. /*?* can also capture useful product variants or language parameters. Robots.txt does not replace canonical signals or automatically resolve duplication. Blocking paginated categories can obstruct discovery of products further down the sequence. Check access to required CSS, JavaScript and images.

Check official definitions for specialist crawlers such as Googlebot-Image and Googlebot-News. AI bots also have different purposes: OpenAI separately manages GPTBot for potential training use and OAI-SearchBot for search. Restricting one does not automatically restrict the other; user-initiated access can behave differently. Consult the official crawler guidance. These preferences are not a universal guarantee against every provider’s data access or use.

06

Add the sitemap and publish a controlled change

A website migration should compare robots.txt settings across old and new environments. Use a complete URL in the Sitemap field; multiple sitemap lines or a sitemap index are possible. Google also supports RSS and Atom feeds, which typically cover recent updates. A sitemap does not guarantee crawling or indexing.

An unrestricted starting example is User-agent: *
Disallow:
Sitemap: https://example.com/sitemap.xml
. example.com is illustrative. Verify the real sitemap and URL goals before copying it, and remember that a specific crawler group can take precedence over this general group.

Record scope and rollback

Save the current content, editing location, critical URL examples and rollback version. Protect staging with access controls.

Make the smallest useful change

Document the reason for each rule and compare test, language, product and resource paths. Include appropriate URLs intended for indexing in the sitemap.

Inspect the published response

Read the file on the live hostname. Do not accidentally carry over staging’s Disallow: / rule; check that caches serve the intended version.

07

Validate actual accessibility after syntax checks

During a technical SEO audit, use the current Search Console robots.txt report to inspect the last fetched version, fetch time and parsing issues. Do not make the retired Robots.txt Tester screen part of your release procedure. A recrawl request after a critical fix does not guarantee immediate fetching of the file or affected pages.

Review ten checks together: lowercase filename; root location; UTF-8; size; crawler groups; sitewide restriction; important content and language paths; required resources; sitemap; published response and rollback owner. For Bing, consult current Webmaster Tools validation. In desktop crawlers such as Screaming Frog, verify the selected user-agent and robots settings. Their output is not proof of Google’s actual crawl.

For an unexpected restriction, identify the line and affected scope, publish the correct version, then verify live access again. Follow up through URL Inspection, indexing reports and server logs. Avoid attributing every traffic change to this file. Record the release date, approver and next check so the issue does not recur during a theme or platform update.

BEFORE YOU DECIDE

Frequently asked questions

Is a missing robots.txt file always a problem?

No. A site without required restrictions may have no file. For Google, a 404 means no valid robots.txt restrictions were found; it does not guarantee that every URL will be crawled or indexed.

Should I Disallow a page with noindex?

Google needs crawl access to read its noindex instruction. A noindex line inside robots.txt is unsupported by Google. Private content also needs appropriate access permissions.

Does Crawl-delay work for Googlebot?

Google does not support that field. Check another crawler’s current documentation before using it. Do not follow instructions for the old Search Console crawl-rate setting; use current crawling and infrastructure guidance for load issues.

How many lines should the file contain?

There is no universal line count. Keep rules understandable; Google ignores content beyond 500 KiB. Controlled path patterns can replace long lists, provided you test URLs that could be captured accidentally.

When will changes take effect?

It depends on fetching and caching. Google usually caches robots.txt for up to 24 hours, potentially longer during errors. Updating the file does not mean that every previously blocked page is immediately recrawled.

Does blocking GPTBot stop all AI access?

No. Crawler purposes and user-initiated access differ. GPTBot and OAI-SearchBot are separate preferences. Robots.txt is neither access security nor a universal mechanism for stopping data use across providers.

LET’S DEFINE THE SCOPE

Review rules against your URL examples

Share the current robots.txt, sitemaps, important product or service paths and crawl issues so we can define the change and its validation steps.

Discuss technical SEO scope

PRIX STUDIO

Let’s talk about your project.

  1. Contact
  2. Project
  3. Review
Let’s get acquainted.
What’s your goal?
Services *Select more than one
Website design
Software development
Mobile apps
Digital advertising
SEO
AI & automation
Design & content
Marketing & growth
One last look.