VantinoBot — crawl policy

VantinoBot is the web crawler operated by Vantino SARL, a Swiss company based in Geneva. It collects public listing pages — events and job postings — so that they can be indexed and shown on our platforms.

This page is the policy our crawler links to from its User-Agent header. If VantinoBot has visited your site and you have a question, an objection, or a removal request, everything you need is below.

We would rather be told to stop than be blocked silently. One email to bots@vantino.com is enough, and we treat it as permanent — see How to block us.

We do not run this crawler for AI training data, we do not resell what we collect, and we do not build advertising profiles.

HOW TO IDENTIFY US

Every request we make carries these headers:

User-Agent: VantinoBot/1.0 (+https://vantino.com/bot)
From:       bots@vantino.com

We use a single identity across our whole fleet. We do not rotate User-Agent strings, use residential proxy pools, or otherwise disguise our traffic to get around a block. If we are blocked, we either respect it, ask you for permission, or drop the source.

There is one narrow exception, and we would rather state it than have you discover it. Some firewalls answer 403 Forbidden to any User-Agent containing the string Bot — including for robots.txt itself. When that happens we retry the robots.txt file only with a neutral library identifier, so that we can read your published rules and obey them. Every content request is always made as VantinoBot/1.0.

WHAT WE CRAWL — AND WHAT WE DO NOT

We fetch public pages that list events or job openings, plus the pages those listings link to. Nothing else.

We do not:

  • solve CAPTCHAs or use CAPTCHA-solving services;
  • bypass paywalls, registration walls or login gates;
  • crawl anything behind authentication;
  • collect personal data such as email addresses or phone numbers from the pages we visit;
  • fetch content marked noindex or nofollow for our user-agent.

Before any source is switched on, a person reviews that site's terms of service and records the review. A source stays off until that has happened — technical access is never our justification on its own.

We also honour noindex, noai and nosnippet directives where they are present.

ROBOTS.TXT

We follow RFC 9309, the robots.txt standard, using the open-source protego parser. We read your robots.txt before fetching anything and re-read it at least every 24 hours.

  • Disallow rules are matched against VantinoBot and the * wildcard before every single request.
  • If robots.txt returns a server error (5xx), we stop crawling that host until it recovers.
  • If it is unreachable or returns 404, we follow the standard and treat the site as crawlable — subject to everything else on this page.

Crawl-delay and Request-rate are honoured as a floor, never a ceiling: they can slow us down, and nothing in our configuration can speed us back up past what you asked for.

Ignoring robots.txt requires a written agreement with the site owner. No source we currently crawl does.

HOW OFTEN WE VISIT

By default we make at most 30 requests per minute per host, with no more than 2 requests in flight at once, spaced with a little jitter so we never arrive in a burst. Most sources are configured well below that.

If you answer 429 Too Many Requests or 503 with a Retry-After header, we pause that host for exactly as long as you asked.

When a request fails we back off progressively — 1 minute, 5, 15, an hour, 6 hours, 24 hours — and after the sixth failure we stop and flag the source for a human. We do not retry indefinitely.

A page that has not changed is not re-fetched: we use ETag and Last-Modified conditional requests, so most repeat visits cost you a 304 Not Modified and nothing else.

HOW TO BLOCK US

Any one of these is enough, so pick whichever is easiest for you.

In your robots.txt — takes effect on our next read, within 24 hours:

User-agent: VantinoBot
Disallow: /

By email — write to bots@vantino.com with the host or URLs you want removed. We action requests within 5 business days and confirm in writing. If it is urgent, put URGENT in the subject line and we will pause the source within hours.

An opt-out is permanent. It does not expire, and we will not come back later to ask you to reconsider.

WHAT WE DO WITH WHAT WE COLLECT

Listings we collect are shown on Vantino platforms with attribution and a link back to your page — we do not republish your content under another identity, and we do not strip the route back to the source.

Raw page copies are kept for 7 days for debugging only. They are never republished and never shared with anyone else. Everything is stored on EU-resident infrastructure; we use no data brokers and no advertising networks.

One third party is involved. When a listing gives an address but no coordinates, we send that address — venue or company name, street, city, country — to the Google Maps Geocoding API to obtain a latitude and longitude. Each address is sent once and the result is cached on our side. Nothing else about the listing, and nothing at all about our visitors, is sent to Google.

CONTACT

Email: bots@vantino.com

We read that inbox and answer it. Questions about how we crawl, requests for removal, or an offer of a proper data feed instead of crawling — all welcome at the same address.

Vantino SARL

PO Box 2224

1211 Geneva 2, Switzerland

General enquiries: info@vantino.com