---
title: Robots.txt: What It Is, Why It Matters, and How to Use It
description: Learn what robots.txt is, why it matters, and how to use it to tell search engines which pages on your site they should not access.
source: https://vitaminaweb.digital/en/blog/robotstxt-what-it-is-why-it-matters-and-how-to-use-it
lang: en
---

![Robots.txt: What It Is, Why It Matters, and How to Use It](https://vitaminaweb.digital/storage/blog/DwjvVlfjlSublqhLHQ5jRul2URi10y7y3XLQJ1CL.webp)

Search engines use bots, known as web crawlers or spiders, to explore the entire internet, indexing most, if not all, of the available content. In response, a standard called the "Robots Exclusion Protocol" was established. It allows you to add a file called robots.txt to your site's root directory, telling search engine bots which pages they should not access.

## Why Robots.txt matters for your website

The robots.txt file is an essential tool for any website project, as it lets search engines know which files or directories they can access. It's important to create a robots.txt file, even if it's empty, in your domain's root directory to ensure search engines can access the entire site if something goes wrong with the server and Google chooses not to read the whole site.

Keep in mind that each site should have only one robots.txt file, and it must be in the root directory. If another robots.txt file is located in any other directory, search engines won't access it. However, this practice may not be ideal for large companies, since not all employees have access to the site's root directory.

Search engine bots are designed to navigate the internet and index content for display in search results. However, there are cases where you may not want certain pages to appear, such as the ones listed below.

**Login pages**
Pages that provide restricted access, such as an intranet, generally shouldn't be indexed;

**Pages with duplicate content**
If you have several landing pages with similar content for your Google AdWords campaigns, it's a good idea to block the copies and index only one version to avoid duplicate content issues;

**Print pages**
If your site has versions for screen and print, it's a good idea to remove the print version from Google's index.

Finally, it's important to note that the robots.txt file is not a security measure. It prevents search bots from reading the specified content, but it does not prevent users from accessing it.

## How to create a robots.txt file

There are several ways to create a robots.txt file, such as using Notepad or any other plain text editor. However, free online tools let you select which pages should be blocked from search bots. The tool provides the complete code, ready to use in your robots.txt file. We recommend trying one of these tools to make creating the file easier.

An example of a [Robots.txt generation tool can be found here](https://seocheckfree.com/pt/robots-txt-generator).

## robots.txt file format and syntax

The robots.txt file syntax is used to create an access policy for bots. It includes reserved words that act as commands to allow or deny access to specific directories or pages on a site. Below are the main robots.txt file commands.

**User-agent**
The User-agent command lists which bots should follow the rules defined in the robots.txt file. For example, if you want only Google's search engine to follow the file's instructions, you need to set the User-agent to Googlebot. Here are the main options:

• Google: User-agent: Googlebot
• Google Images: User-agent: Googlebot-images
• Google AdWords: User-agent: Adsbot-Google
• Google AdSense: User-agent: Mediapartners-Google
• Yahoo: User-agent: Slurp
• Bing: User-agent: Bingbot

All search engines: User-agent: \* (or simply leave out the User-agent command)

**Sitemap**The Sitemap command lets you specify the path and name of the site's XML sitemap. However, Google's Webmaster Tools offers greater control and visibility for this function. See how Google lists several sitemaps in its robots.txt file:

• Sitemap: http://www.google.com/hostednews/sitemap_index.xml
• Sitemap: http://www.google.com/sitemaps_webmasters.xml
• Sitemap: http://www.google.com/ventures/sitemap_ventures.xml
• Sitemap: http://www.gstatic.com/dictionary/static/sitemaps/sitemap_index.xml
• Sitemap: http://www.gstatic.com/earth/gallery/sitemaps/sitemap.xml
• Sitemap: http://www.gstatic.com/s2/sitemaps/profiles-sitemap.xml
• Sitemap: http://www.gstatic.com/trends/websites/sitemaps/sitemapindex.xml

**Disallow**
The Disallow command tells search engines which directories or pages should not be included in the index. Here are a few examples:

• Disallow: /prod - tells robots not to index folders or files that start with "prod";
• Disallow: /prod/ - tells robots not to index content inside the "prod" folder;
• Disallow: print1.html - tells robots not to index the content on the print1.html page.

**Allow**
The Allow command tells robots which directory or page should have its content indexed. Directories and pages are allowed by default, so this command should only be used when a webmaster has blocked access to a directory with the Disallow command but wants a file or subdirectory within the blocked directory to be indexed. See the example below:

• Disallow: /catalogs
• Allow: /catalogs/about

Allow lets the /about directory under the /catalogs directory be indexed.

## Final tips for getting the most out of your robots.txt

**Understand the syntax**
Before creating your robots.txt file, it's essential to understand its syntax so you can avoid mistakes and create the right rules.

**Use the right keywords**
To specify which bots should follow the rules, which pages should be blocked, and which pages should be allowed, use the appropriate keywords, such as "User-agent", "Disallow", and "Allow".

**Don't rely entirely on robots.txt**
Keep in mind that the robots.txt file is only a suggestion for search bots, and some may ignore the rules or interpret them differently.

**Use testing tools**
Before implementing the robots.txt file on your site, use testing tools to check that the rules are working correctly and that the pages you want to block are actually being blocked.

**Update it regularly**
It's important to keep your robots.txt file up to date, especially when making changes to your site that could affect indexing. If you remove a page or directory from your site, remember to remove the corresponding Disallow rule from the robots.txt file.

**Be careful when using Disallow**
Use the Disallow command with caution, as it can block important pages on your site, such as product or category pages. Make sure the pages you block aren't essential for indexing and the user experience.

**Include a sitemap**
Add an XML sitemap to your site to help search bots find all its important pages, which can help ensure they are all indexed correctly.

## Read also

     What's New in Digital

### [Instituto da Maturidade Digital launches Agency Directory portal, developed by VitaminaWeb](https://vitaminaweb.digital/en/blog/instituto-da-maturidade-digital-launches-agency-directory-portal-developed-by-vitaminaweb)

Instituto da Maturidade Digital launches Agency Directory, a platform developed by VitaminaWeb to organize and make it easier to discover compani...

 03 Oct, 2026 · 5 min      Development

### [All-in-One WP Migration flaw could put millions of WordPress sites at risk](https://vitaminaweb.digital/en/blog/all-in-one-wp-migration-flaw-could-put-millions-of-wordpress-sites-at-risk)

CVE-2026-19949 affects versions up to 7.109 of the popular backup and migration plugin. An attack could escalate from SQL injection to remote cod...

 03 Sep, 2026 · 5 min      Development

### [Web Application Security: A Strategic Guide from Vulnerability to Digital Maturity](https://vitaminaweb.digital/en/blog/web-application-security-a-strategic-guide-from-vulnerability-to-digital-maturity)

This guide offers an in-depth strategic analysis of the main risks, best mitigation practices, and the integration of security into the software...

 23 Feb, 2026 · 5 min
