TarkaBot
TarkaBot is the crawler of Tarka, a web search engine being built in India. If it has visited your site, this page says what it was doing and how to control it.
How to recognise it
Every request carries this user agent and a contact address in the From header:
Mozilla/5.0 (compatible; TarkaBot/0.1; +https://gettarka.com/bot) From: bot@gettarka.com
What it does
- Reads your
robots.txtbefore anything else and follows it. If the file cannot be read because of a server error, it leaves the site alone. - Makes one request at a time to a site, at least a second apart, and slows further if your server is slow to answer.
- Backs off when it gets a
429or503, and waits as long asRetry-Aftersays. - Fetches HTML pages only. It does not run JavaScript, submit forms, or log in.
- Follows
noindexandnofollow, whether in a<meta name="robots">tag, anX-Robots-Tagheader, or a link'srel.
How to control it
To keep TarkaBot away from part of your site, or all of it, add a group for it to robots.txt:
User-agent: TarkaBot Disallow: /private/
User-agent: TarkaBot Disallow: /
To slow it down, set a delay in seconds between requests:
User-agent: TarkaBot Crawl-delay: 10
Rules for User-agent: * apply to TarkaBot when there is no group naming it. Changes are picked up within 24 hours.
How to reach us
- Questions about the crawler, requests to slow down or stop, and requests to remove pages: bot@gettarka.com
- Reports of network abuse: abuse@gettarka.com
Please include your domain and, if you can, a few lines from your server log. We read both inboxes.