Free tool by MD Niamul

Free AI Crawler & robots.txt Checker See Which AI Bots Can Read Your Website

Find Out If ChatGPT, Claude, Perplexity and Google’s AI Can Find and Cite You

No login needed
26 bots tested
Free Excel report

Enter your website and this free AI crawler checker reads your robots.txt, your llms.txt and your homepage tags, then tells you which AI bots are allowed or blocked: GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, Google-Extended and more. It shows the exact line that decides each result and includes a robots.txt AI bots rule builder and an llms.txt generator.

26Bots tested
4Policy presets
850+Projects delivered
600+Five-star reviews
MD Niamul, creator of the free ai crawler & robots.txt checker
upworkTop Rated Plus
AI Search + robots.txtFree Checker Tool
OfficialStape Partner

Everything the AI Crawler Checker Tests

GPTBot
OAI-SearchBot
ChatGPT-User
ClaudeBot
Claude-SearchBot
PerplexityBot
Google-Extended
Googlebot
Bingbot
Applebot-Extended
meta-externalagent
Amazonbot
CCBot
robots.txt Parser
Sitemap Lines
llms.txt Check
llms.txt Generator
Excel Export
FREE AI CRAWLER CHECK

Check Which AI Bots Can Read Your Site

Enter your website. Add a page path if you want to test one page or folder as well as the homepage.

AI Crawler & robots.txt Checkerrobots.txt · AI search bots · Training bots · llms.txt · Meta robots
Free · No login

The address you enter is sent to our server, which loads your robots.txt, llms.txt and homepage once each. We keep a simple log of the address checked and the result. No personal details are saved unless you ask for the report by email.

Shows the line that decides each botBuilds the rules to addExcel report to download or email
WHAT IT CHECKS

What Does the AI Crawler Checker Test?

Three public files, read once, then tested against the rules crawlers actually follow.

AI Search Bots

OAI-SearchBot, Claude-SearchBot, PerplexityBot and others that let AI products find and cite your pages.

AI Training Bots

GPTBot, ClaudeBot, Google-Extended, CCBot and others that collect pages for model training.

User-Triggered Fetchers

ChatGPT-User, Claude-User and Perplexity-User, which open a page when a person asks about it.

robots.txt Rules

Every group and rule, with the line that decides the answer for each bot.

One Page or Folder

Test any path, not only the homepage, including * and $ patterns.

Homepage Tags

Meta robots and X-Robots-Tag: noindex, nosnippet, max-snippet and noai.

llms.txt

Whether the file exists and follows the proposed format: title, summary and link sections.

Rule and File Builders

A robots.txt rule builder with four policies and an llms.txt generator filled from your homepage.

How Does the AI Crawler Checker Work?

Three steps, usually a few seconds.

01

Enter your website

Add a page path too if you want to test a specific page or folder.

02

The files are read and tested

Your robots.txt is parsed the way crawlers parse it, then each of the 26 bots is matched against it.

03

See who is allowed, then fix it

Pick a policy and copy the lines to add to robots.txt. Generate an llms.txt if you want one.

What Is an AI Crawler?

An AI crawler is a program that reads web pages for an AI company. Some collect pages to train models. Others build a search index so an assistant can find, quote and link to your page. A third kind fetches one page at the moment a person asks about it. Each has its own name, called a user-agent, and your robots.txt file can give each name different rules.

Why does it matter for a business?

  • Being found: if the search bots of ChatGPT, Claude or Perplexity are blocked, those products cannot cite your pages from their own crawl.
  • Accidental blocks: many “block AI” lists mix training bots and search bots. Sites copy them and disappear from AI search without meaning to.
  • Your content, your choice: you may want to be cited but not used for training. robots.txt lets you say so, bot by bot.

How are robots.txt rules matched?

This checker follows the published robots.txt standard (RFC 9309) and Google’s documentation:

  • A crawler uses the group that names it. Names are not case-sensitive. If no group names it, it uses the User-agent: * group.
  • Several groups for the same crawler are combined.
  • The longest matching path wins. If an Allow and a Disallow are the same length, Allow wins.
  • * matches any text and $ marks the end of the address.
  • An empty Disallow: allows everything. Unknown lines are ignored.
  • A bot with its own group ignores the * group. This catches many people out.

What is llms.txt?

llms.txt is a proposed Markdown file at /llms.txt that lists your most useful pages for AI tools. The proposal asks for one H1 title, a short summary in a blockquote, then H2 sections with links. It is a proposed convention, not a standard, and the large AI companies have not committed to reading it. It is cheap to add, so the tool includes an llms.txt generator, but do not expect it to change rankings.

What this tool cannot check

  • robots.txt is a request, not a wall. Well-behaved bots follow it. Some user-triggered fetchers say openly that they may not.
  • It cannot see firewall or CDN bot blocking, such as Cloudflare’s AI bot blocking, or rate limits. A bot can be allowed in robots.txt and still be stopped there.
  • It cannot tell you whether an AI product actually cites you. Being allowed does not guarantee being crawled.
  • Homepage tags are read from the homepage only, not every page.
  • The bot list reflects each company’s published documentation when this tool was built. Companies add and rename bots. ByteDance and Cohere publish little about theirs, so treat those two rows as a guide.

AI Bots and What They Do

The main user-agent names, from each company’s own documentation.

CompanyTrains modelsPowers AI searchFetches for a user
OpenAIGPTBotOAI-SearchBotChatGPT-User
AnthropicClaudeBotClaude-SearchBotClaude-User
Perplexity–PerplexityBotPerplexity-User
GoogleGoogle-Extended (control name)Googlebot–
AppleApplebot-Extended (control name)Applebot–
Metameta-externalagentmeta-webindexermeta-externalfetcher
AmazonAmazonbotAmzn-SearchBotAmzn-User
MistralMistralAI-TrainingMistralAI-IndexMistralAI-User
Common CrawlCCBot (open archive)––

Who Is This Checker For?

Business owners

See in plain words whether AI assistants can read your site.

SEO specialists

Audit AI access and catch copied block lists that hide a site from AI search.

Publishers

Allow citation while opting out of model training.

Developers

Test a path against wildcard rules before you ship a robots.txt change.

MORE FREE TOOLS

More Free Marketing Tools by MD Niamul

33 free tools for tracking, analytics, advertising and SEO. No login needed.

Free Social Share & SERP Preview Checker

Preview how a page looks on Google, Facebook, LinkedIn and X, and fix its Open Graph tags.

Open the preview checker →

Free Shopify Store Checker

See a Shopify store’s theme, apps, catalogue size, tracking and speed.

Open the Shopify checker →

Free SPF, DKIM & DMARC Checker

Check the DNS records that decide whether your marketing emails reach the inbox.

Open the email checker →

Free CMS Checker

Find out which CMS, theme, plugins and technology any website is built with.

Open the CMS checker →

Trusted by 600+ Brands & Agencies

MYLENE
HOVER BOARDS
GT SOLI
CAMBRIDGE
BIENEN CORB 24
ZULU SHACK CREATIVE
ADPSY LLC
HIGH PERFORMANCE MEDIA
ADS PERFORMANCE
LEGESI
TESTIMONIAL

Analytics & Conversion Tracking Client Testimonials

Hear from our clients. Loved by 600+ businesses worldwide.

Madison Jonas, client of MD Niamul
Madison JonasGoogle Ads clientLinkedIn recommendation
★★★★★

“I had an excellent experience working with MD on my Google Ads account. Everything was set up perfectly and worked smoothly without any issues or troubleshooting needed. His expertise and attention to detail made the whole process effortless.”

Mehboob Khan, client of MD Niamul
Mehboob KhanGoogle Analytics projectsLinkedIn recommendation
★★★★★

“I worked with Niamul on multiple Google Analytics projects and was impressed by his expertise and precision. He has a strong grasp of tracking, reporting, and optimization, always ensuring accurate insights. He is proactive, reliable, and easy to collaborate with.”

Jacco Bouw, client of MD Niamul
Jacco BouwShopify, GA4 & Google Ads trackingLinkedIn recommendation
★★★★★

“MD Niamul is extremely skilled. He fixed my Shopify, Google Analytics, and Google Ads conversion tracking perfectly. Everything works exactly as it should now, and he even added an extra data layer, which was very helpful. Outstanding service, fast delivery, and highly recommended.”

Ryan S Kemp, client of MD Niamul
Ryan S KempCMO, Co-Founder, Zulu Shack CreativeLinkedIn recommendation
★★★★★

“Reliable, knowledgeable, and truly trustworthy. I’ve been working with Niamul for over a year, and he consistently delivers exceptional results across all analytics tasks. His expertise, communication, and commitment make him my go-to specialist for tracking, measurement, and data accuracy.”

Casper Klein, client of MD Niamul
Casper KleinMeta Pixel & server-side trackingLinkedIn recommendation
★★★★★

“MD N delivered flawless Meta Pixel, CAPI, and server-side tracking. He understood our goals quickly, explained everything clearly, and ensured full transparency. Communication was smooth, updates were consistent, and the results improved our data accuracy and ad performance.”

600+

Five-star reviews from businesses and agencies in the USA, Canada, UK and Australia.

Read All Reviews

Video Testimonials: Hear It From Our Clients

FAQ

Frequently Asked Questions About the Free AI Crawler & robots.txt Checker

Is the AI crawler checker free?

Yes. It is free, needs no login and tests 26 bots in one check. You can download the result as an Excel file.

How do I know if ChatGPT can read my website?

Enter your site. The tool tests OAI-SearchBot, which powers ChatGPT search, GPTBot, which collects training data, and ChatGPT-User, which opens a page when a user asks. Each is shown as allowed, blocked or partly blocked.

What is the difference between GPTBot and OAI-SearchBot?

GPTBot collects pages that may be used to train OpenAI’s models. OAI-SearchBot is used to show websites in ChatGPT’s search results. You can block one and allow the other.

Does blocking Google-Extended remove me from Google Search?

No. Google says Google-Extended does not affect inclusion or ranking in Google Search. It controls whether your content is used to train Gemini models and for grounding in Gemini.

Can I allow AI search but block AI training?

Yes. Use the “Allow AI search, block AI training” preset in the Fix / generate tab. It writes the lines to add to your robots.txt.

Will a blocked bot really stay away?

Well-behaved bots follow robots.txt, but it is a request, not a wall. Some user-triggered fetchers say they may not follow it. To enforce a block you need a firewall or CDN rule.

What is llms.txt and do I need one?

It is a proposed text file that lists your key pages for AI tools. It is not an agreed standard and the large AI companies have not committed to reading it. It is optional and costs little to add.

Why does the tool say a bot is allowed when I have no rule for it?

A bot with no group of its own follows the “User-agent: *” group. If nothing there blocks the page, the bot is allowed. With no robots.txt at all, everything is allowed.

Can this tool tell me if AI assistants cite my site?

No. It checks whether you allow them to read your site. It cannot see firewall blocks or whether any AI product has crawled or cited you.

Know Who Reads Your Site, and Measure What It Brings

I am MD Niamul, a conversion tracking and web analytics specialist and an official Stape partner. I set up tracking that separates visits and leads from AI assistants, search and ads, so you can see what each one is worth.