The DataHarmoni scanner
If you saw this name in your server logs, this page explains who we are, what it does, and how to stop us.
What it is
The scanner is a small program we run to check public web pages for common accessibility problems (for example missing alt text, low colour contrast and unlabelled links). It is run by DataHarmoni, B-13, Anupam Enclave Phase 1, New Delhi 110068, India. It identifies itself with this user agent:
DataHarmoniScanner/0.2 (+https://dataharmoni.com/scanner)
What it does and does not do
The scanner only requests pages that are publicly accessible without authentication or circumvention. It does not attempt to bypass access controls, exploit vulnerabilities, submit forms, or access non-public resources.
- It requests up to five public pages of a site, about two seconds apart, one at a time.
- It never logs in, never submits forms and never tries to get past a login or a paywall.
- It only connects to public internet addresses. It does not visit internal or private addresses.
- It runs in a normal browser, so your pages may count it as a visit.
- It starts at the home page and follows links from it. It does not guess other addresses or probe for vulnerabilities.
- If a site owner has asked us to audit their own store, that is a separate paid service done with the owner's permission. In that service the scanner may also request the standard addresses of the cart and sign-in pages, and may use access the owner gives us. The free snapshot and our own research scans never do.
A second small tool
To find the business email address a company publishes, a second tool called DataHarmoniProspector reads the company's robots.txt, its home page and, if the home page links one, its contact page: at most three requests, two seconds apart. It follows the same rules as this page and respects a robots.txt rule for DataHarmoniProspector. It only records email addresses of a role kind (for example info@ or hello@) that the company itself publishes.
robots.txt
Before scanning, it reads your robots.txt. A rule for DataHarmoniScanner is followed; if there is none, the rules for all crawlers (User-agent: *) are followed. Robots.txt is one signal we use to decide whether to visit a page. It is not a licence to access your site, and it does not replace any terms or access controls of yours; if you ask us to stop by any route, we stop.
What we do with the result
We use it to prepare a one-page snapshot for the site owner. See why you may have received an email from us and our privacy notice.
Opening up more of your site
Most of the time, nothing. Once you've told us to go ahead with an audit, we test the pages agreed with you ourselves — you don't need to change anything. Two exceptions:
- You'd rather grant it directly in your own
robots.txt(for example, if a page like your cart is normally disallowed and you want the permission to be visible and standing, not a one-off): add a rule naming our scanner specifically, so nothing else you block is affected. For example:User-agent: DataHarmoniScanner, thenAllow: /cart/. - Your site sits behind bot-protection or a WAF (Cloudflare, Akamai and similar).
robots.txtalone won't reach us in that case; you'd add an allow-rule for our user agent in that service's own console. Email hello@dataharmoni.com if you'd like help with this.
How to opt out
Any of these works:
- Add this to your
robots.txt:User-agent: DataHarmoniScannerDisallow: /
and, for the second tool, the same withUser-agent: DataHarmoniProspector - Email hello@dataharmoni.com with your domain. We add it to a block list the same working day and confirm within two working days.
- Block the user agent shown above.
If the scanner causes a problem on your site, email us and we will stop it straight away.