# For AI agents

URL: https://igregulator.io/docs/for-ai-agents/
Markdown: https://igregulator.io/docs/for-ai-agents.md

> Build AI agents on iGregulator: llms.txt, markdown docs, the MCP server, structured errors and rules for phrasing a licence verdict safely. No key to start.

Building an AI agent, LLM-powered integration, or automated compliance
tooling? Everything you need to wire iGregulator into a language-model
workflow lives here.

## Quick discovery

Three machine-readable resources document the entire product. One
fetch each — no crawling required.

- **[llms.txt](https://igregulator.io/llms.txt)** — structured index per the
  [llmstxt.org spec](https://llmstxt.org). ~15 KB, complete product
  map with links into the longer docs.
- **[llms-full.txt](https://igregulator.io/llms-full.txt)** — every `/docs/*` page
  concatenated into one file. ~260 KB. One fetch buys full context.
- **[OpenAPI 3.1 spec](https://api.igregulator.io/openapi.json)** —
  authoritative API schema. Every endpoint, request/response shape,
  auth scheme, rate-limit description.

Every page on this site also emits `<link rel="alternate">` pointers
to all three URLs so a crawler hitting the landing can auto-discover
them. The homepage additionally returns `Link` response headers
(RFC 8288 / RFC 9727) pointing at the catalog, OpenAPI spec, docs, and
server card — so headless agents find them without parsing HTML.

### Well-known discovery endpoints

Standards-based entrypoints under `/.well-known/` (and the apex), for
agents that look there first:

- **[/.well-known/mcp/server-card.json](https://igregulator.io/.well-known/mcp/server-card.json)**
  — MCP Server Card (SEP-1649): server info, transport endpoint, and the
  full tool list. (Legacy [/.well-known/mcp.json](https://igregulator.io/.well-known/mcp.json)
  is still served too.)
- **[/.well-known/api-catalog](https://igregulator.io/.well-known/api-catalog)** — API catalog
  (RFC 9727) linking the OpenAPI spec, docs, `llms.txt`, and the
  `/v1/health` status endpoint.
- **[/.well-known/agent-skills/index.json](https://igregulator.io/.well-known/agent-skills/index.json)**
  — Agent Skills discovery index (v0.2.0). Currently ships a
  `verify-gambling-license` skill (digest-pinned `SKILL.md`).
- **[/auth.md](https://igregulator.io/auth.md)** — how to authenticate: bearer API keys for the REST
  API and MCP server, or OAuth 2.1 sign-in for MCP connectors at
  `https://mcp.igregulator.io/mcp/account` (its RFC 9728 metadata is on
  `mcp.igregulator.io`; the authorization server is `app.igregulator.io`).

Content usage is declared via `Content-Signal` in
[robots.txt](https://igregulator.io/robots.txt) — `search`, `ai-input`, and `ai-train` are all
permitted.

## Agent-friendly features

Deliberate design choices that make integration cleaner for agents.

### Structured error details

Every error response carries `details.reason` + `details.suggestion`.
Agents branch on `reason` instead of free-form text and surface
`suggestion` to the user / caller as-is.

```json
{
  "error": "domain is not a valid hostname",
  "code": "invalid_query",
  "details": {
    "field": "domain",
    "reason": "not_a_valid_hostname",
    "suggestion": "Pass a bare hostname — no scheme, no path, no underscores. Example: 'paddypower.com' or 'www.bet365.com'."
  }
}
```

The full `reason` vocabulary is stable; see
[error handling](https://igregulator.io/docs/errors/) for the code + reason matrix — branch on
`code` **and** `reason`, because one `code` (`rate_limited`,
`quota_exceeded`, `invalid_query`) covers several causes.

### `verdict` — read this first

`/v1/check` answers with a top-level `verdict` — `licensed`,
`licensed_provisional`, `licence_not_active`, `domain_not_listed`,
`related_host_listed`, `name_match_only`, `not_found` or `generic_term` —
derived from `status`, `domain_status`, `status_qualifier` and `confidence` by
fixed rules ([confidence scoring](https://igregulator.io/docs/confidence/#verdict--the-answer-in-one-field)),
so an agent does not have to combine them itself. `status: active` with
`domain_status: delisted` is `domain_not_listed`, not "licensed"; a revoked or
suspended licence is `licence_not_active` even when the domain is de-listed too.
Only `licensed` and `licensed_provisional` mean licensed now — treat every other
value, including one you have not seen before, as "not confirmed as licensed".

`verdict_detail` is one plain-English sentence to quote as it stands: it names the
regulator, operator and licence (never our KH/TGC/IOM reference as the
regulator's number), dates the read that listed the domain and names any register
behind the answer whose last read is past its freshness window, scopes a miss to
the registers we cover, and never calls anything "unlicensed". Branch on
`verdict`; quote `verdict_detail`.

`www.example.com` and `example.com` are the same site: either finds the host the
register lists (`match.matched_domain` says which). Any other subdomain is a
different host. When the host you asked about is not listed but others on the
same registrable domain are, the answer is never `licensed`: it is
`related_host_listed` (or `licence_not_active` / `domain_not_listed` when that
other host's licence or listing is not in force), and `related_hosts[]` names
those hosts, each with its operator. Say which hosts are listed and that this one
is not; never carry their licence over to it.

### Response `_meta` field

Data-returning endpoints include a `_meta` envelope with provenance:

- `scraped_at` — ISO-8601 timestamp when we last pulled this record.
- `source_url` — exact regulator URL, or null if the row doesn't map to
  one URL. For a licence this is the page that published its **status**
  (`status_source_url`) — an enforcement register or a revoked-licences
  list, say — which is often not the register the licence is listed in.
- `confidence_hint` — `authoritative` (direct register dump),
  `scraped` (HTML / PDF), or `derived` (fuzzy match, not a direct
  lookup).
- `source_modified_at` — reserved for the regulator's own
  modification timestamp. Always `null` today: we don't store it for
  any register yet.

On `/v1/check`, `_meta` also carries:

- `checked_at` — when the answer was computed. (On a name match or a
  miss, `scraped_at` is that same request time — there is no record
  behind the answer.)
- `register` — `{ jurisdiction, last_read_at, fresh, sla_hours }` for the
  matched jurisdiction: when we last read its register and whether that is
  inside its freshness SLA (the same rule as `/v1/health/coverage`).
- `stale_jurisdictions` — on a miss, a name-only match or a related-host answer,
  `[{ code, last_read_at }]` for every covered register past its SLA. A
  site listed there since our last read is not in our data yet: say so
  next to a "not found".

And `match` carries what a citation needs without a second call:
`regulator_name`, `license_id` (for `GET /v1/licenses/{id}`),
`license_reference_is_ours`, `status_source_url` and `status_observed_at`
(where this status was published, and the latest read that still said it),
`matched_domain` (the host the register lists — `www.x.com` for `x.com`) and
`domain_last_listed_at` (when a regulator source last listed it; `null` once
de-listed).

An agent can say "verified via UKGC official register, scraped 6 h
ago" with a real evidence trail, not a gloss. Which page can set which
status in each jurisdiction, and what we keep as evidence, is on the
[methodology](https://igregulator.io/methodology/) page.

### `verification_url` — cite the regulator, not us

Where the regulator runs its own per-domain verification page —
Curaçao's certificate portal, Tobique's validation seal —
`match.verification_url` on `/v1/check` links straight to it. That's
a primary-source citation an agent can hand to the user ("verify on
the regulator's own page"), and it's how we keep the row fresh: every
stored verification page is re-read on a rolling nightly pass (every
Tobique seal each night, each Curaçao certificate every 2–3 days), and
`domain_status` moves
only on a clean read of the regulator's own words. `null` for
regulators with no such surface (UKGC, MGA, KH, AN, IOM). A Curaçao link stored
on the legacy host `cert.gcb.cw` is served on `cert.cga.cw`.

**Before you cite it, read `verification_page_status`** — what that page printed
about the licence at our latest read, verbatim (`Active`, `Revoked`, a Tobique
seal's `VALID`), with `verification_page_read_at`. The same pair is on every
`jurisdictions[]` row and on each `domains[]` entry of `/v1/operators/{slug}`. If
the word does not say the licence is in force, the page is not proof: say what it
reads and when, and don't offer it as confirmation. It changes no status on its
own — a revocation comes only from a regulator publication (`status_source_url`).

### `jurisdictions[]` — dual-licensed domains

A brand can be licensed by different legal entities in different
jurisdictions at once. When a matched domain has more than one
(operator, jurisdiction) pair, `/v1/check` carries a best-first
`jurisdictions[]` array with per-link `status`, `status_qualifier`,
`domain_status`, `license_reference_is_ours`, `verification_url` and
`register` (when we last read that link's register, and whether the read is
fresh). **Report all of them** — the same domain can be active under one
register and delisted under another, and naming only one is a wrong answer. See
[confidence scoring](https://igregulator.io/docs/confidence/#jurisdictions--dual-licensed-domains).

### `/v1/health/coverage`

Public endpoint exposing per-jurisdiction scraper freshness.
`status: healthy` vs `degraded`, `age_hours`, `record_count` per
regulator. See [endpoints](https://igregulator.io/docs/endpoints/). Useful for SLA
dashboards and status pages.

### Stable `operationId`s

Every endpoint has a stable `operationId` (e.g. `checkDomain`,
`checkDomainBatch`, `searchOperators`, `getOperator`,
`getOperatorRegulatoryActions`, `listJurisdictions`, `getLicense`,
`getLicenseHistory`, `checkCoverage`). SDK generators and MCP
servers use these as function names — no `getV1CheckDomain` slug
noise.

### Dual rate-limit headers (keyless calls)

A call made **without a key** to a public endpoint carries the policy in
both a custom and the IETF-draft format. Parse either:

```
X-RateLimit-Policy: tier=public;limit=10;window=hour
RateLimit-Policy: "default";q=10;w=3600
```

The IETF draft (`draft-ietf-httpapi-ratelimit-headers`) is what
Cloudflare, Kong, and similar gateways auto-parse; the custom format
is human-readable for logs.

Keyed calls don't send `RateLimit-Policy`. Their `X-RateLimit-Limit` /
`-Remaining` / `-Reset` describe your **monthly** quota, and
`X-RateLimit-Policy` appears only as `tier=unlimited` (no monthly cap) or
`tier=authenticated` (a key on a public endpoint). See
[rate limits](https://igregulator.io/docs/rate-limits/#response-headers).

### Deprecation lifecycle

A field marked for removal is announced with standard headers (example
shape):

```
Deprecation: @1790812800
Sunset: Mon, 01 Mar 2027 00:00:00 GMT
Link: <https://igregulator.io/docs/changelog/>; rel="deprecation"; type="text/html"
```

`Deprecation` is [RFC 9745](https://www.rfc-editor.org/rfc/rfc9745): a
structured-field date (`@` + Unix epoch seconds) saying when the field was
deprecated, plus the `rel="deprecation"` link. `Sunset` is
[RFC 8594](https://www.rfc-editor.org/rfc/rfc8594): an HTTP-date for the
removal, minimum 90 days out. An agent caching request shapes can inspect
`Sunset` before assuming stability. No fields are currently deprecated,
so the API sends neither header today.

## MCP server

Live at [`mcp.igregulator.io`](https://mcp.igregulator.io). Streamable
HTTP transport (current MCP spec — SSE used for streaming responses).
Compatible with Claude Desktop, Cursor, Windsurf, Cline, and any other
client that speaks MCP.

Tools exposed (lean output — a compact verdict, not the full REST payload):

- `check_domain` — verify a licence by domain (supports `as_of`)
- `check_domain_batch` — up to 100 domains in one call (KYB sweep)
- `search_operators` — search the register by name
- `get_operator` — full operator detail (supports `as_of`)
- `get_operator_regulatory_actions` — enforcement history (fines, suspensions)
- `check_coverage` — data freshness per jurisdiction
- `list_jurisdictions` — all covered regulators
- `get_jurisdiction` — single regulator metadata
- `get_license` — single licence detail
- `get_license_history` — status-change timeline

`check_domain` and every `check_domain_batch` row lead with the API's `verdict`
and `verdict_detail` — branch on the first, quote the second. On a no-match,
`check_domain` also returns `match_absence_reason` + `checked_jurisdictions` —
never collapse a miss into a bare "unlicensed".
On a match, `verification_url` (the regulator's own page) and — for
dual-licensed domains — a compact `jurisdictions[]` array pass through,
so the verdict can be cited and never flattened to one register.

Auth = same API keys as the direct HTTP API — and optional: with no
`Authorization` header, `check_domain`, `search_operators` (top 3),
`list_jurisdictions` and `check_coverage` answer under the REST API's
public limit of 10 requests per hour per IP; the other six tools need a
key. Each keyless tool call counts against the per-IP counter of the REST
route it calls, from your IP: `check_domain` spends the same `/v1/check`
allowance as a direct request ([rate limits](https://igregulator.io/docs/rate-limits/#public-endpoints-no-key)). Setup walkthrough +
example prompts at [/docs/mcp](https://igregulator.io/docs/mcp/). Discovery via the
[MCP Server Card](https://igregulator.io/.well-known/mcp/server-card.json) (SEP-1649) and the
legacy [/.well-known/mcp.json](https://igregulator.io/.well-known/mcp.json).

## WebMCP (in-browser tools)

Distinct from the server above: the homepage registers
[WebMCP](https://webmcp.org/) tools via `navigator.modelContext`, so an
agent **driving a browser** can act without wiring up the HTTP/MCP
integration at all. Two read-only tools, backed by the public API
(10 req/IP/hour, no key):

- `check_gambling_license` — verify by domain or licence number.
- `search_gambling_operators` — search operators by name.

They're feature-detected, so they simply don't appear in browsers
without the WebMCP API. Use the server-side MCP for production
integrations; WebMCP is the zero-setup path for browser agents.

## Integration patterns

### Merchant onboarding verification

Payment-processor KYB flow:

1. User submits a merchant application with a domain.
2. Agent calls `GET /v1/check?domain=X`.
3. Branch on `verdict`, and put `verdict_detail` in front of whoever decides:
   - `licensed` → approve. If a `jurisdictions[]` array is present, the domain is
     on more than one (operator, jurisdiction) licence and each entry has its own
     `status` and `domain_status` — read them all and report the full picture, not
     just `match` ([confidence scoring](https://igregulator.io/docs/confidence/)).
   - `licensed_provisional` (`status: active` with `status_qualifier:
     provisional_under_assessment`) → licensed now, provisionally: a Curaçao
     licence past its stated term whose final assessment by the CGA is outstanding
     (the CGA keeps it in force until it decides). Approve if your policy accepts
     provisional licences, say so in the answer, and re-check — it can become
     final or end.
   - `licence_not_active` → **do not auto-approve.** `verdict_detail` names the
     status. `revoked` / `suspended` mean a regulator published an enforcement
     decision against that licence — cite `match.status_source_url`. `expired`,
     `surrendered` (the operator gave it up) and `not_in_register` (the register
     no longer lists it, and published no reason) mean "not licensed right now",
     and `unknown` means wording we could not classify — none of them is a
     revocation, so route them to a human rather than rejecting with an
     enforcement-flavoured reason. `upstream_status` is the regulator's own word;
     for since when a licence has been unlisted, fetch `match.license_id` from
     `GET /v1/licenses/{id}` (`not_listed_since`, `last_listed_at`).
   - `domain_not_listed` → **do not auto-approve**: the licence is active, but the
     regulator no longer lists this domain on it.
   - `related_host_listed` → **manual review**. This host is not listed; other
     hosts on the same registrable domain are (`related_hosts[]`, each with its
     operator). The applicant's site may be one of them under another name, or a
     different site altogether — ask, then check the host they confirm.
   - `name_match_only` → manual review. The domain is on no licence we read; a
     name matched an operator (`confidence: medium` is a close match, `low` a weak
     one), so `status` is that operator's licence, not this site's. Never reject
     on it — and never approve on it either.
   - `generic_term` → manual review. The domain root is a generic gambling label
     (`casino.org`, `poker.com`) and `/v1/check` refuses to guess which operator
     runs it. The site may well be licensed, just not identifiable from the label.
   - `not_found` → not in any covered register. Scope it to
     `checked_jurisdictions` ("not found in the registers of the 7 jurisdictions we
     cover"), never an unqualified "unlicensed", and mention any
     `_meta.stale_jurisdictions` — `verdict_detail` already does both.
   - A value not listed here → treat it as "not confirmed": manual review.

### Daily regulatory sweep

Portfolio-monitoring automation:

1. Agent keeps a list of N operator slugs in its CRM / knowledge base.
2. Daily cron: iterate the list, call `GET /v1/operators/:slug` (licences
   + domains) and `GET /v1/operators/:slug/regulatory-actions`
   (enforcement history — not included in the operator detail). An empty
   list there means no action is *linked* to that operator, not a clean
   record: many published actions aren't matched to an operator (UKGC
   107 of 107 linked, MGA 2 of 160, CW 0 of 18 on 2026-09-28). The response
   says so itself: `_meta.note` (quote it on an empty list) and
   `_meta.sources_read`, the regulator publications we read.
3. Diff against previous day's snapshot. Detect status changes,
   regulatory actions, expiry windows. (Or let
   [webhooks](https://igregulator.io/docs/webhooks/) on a [watchlist](https://igregulator.io/docs/watchlist/) push
   them to you.)
4. Alert compliance team on anomalies.

### Domain reputation scoring

Risk-scoring for gambling-adjacent domains:

1. Agent receives an unknown gambling-related domain.
2. Calls `GET /v1/check?domain=X`.
3. When the verdict ties the domain to a licence (`licensed`,
   `licensed_provisional`, `licence_not_active`, `domain_not_listed`), fetches
   the operator's enforcement history from
   `GET /v1/operators/{slug}/regulatory-actions` and combines `verdict` + any
   regulatory actions into a composite risk score. A `name_match_only` operator is
   not tied to the domain: its record says nothing about this site.
4. Score feeds the downstream decision (list / delist / require
   extra verification).

### Bulk KYB sweep (batch)

Auditing a whole merchant book or affiliate list in one shot:

1. Collect the domains (chunks of 100).
2. `POST /v1/check/batch` with `{ "domains": [...] }` — one tool-call per
   100 instead of one per domain.
3. Iterate `results`: each row carries `verdict` + `verdict_detail` +
   `match` + `confidence` + `match_absence_reason` (same semantics as the
   single check), plus `related_hosts[]` when the host isn't stored but others
   on its registrable domain are, and `query.input`, the string you sent. A row has no `jurisdictions[]` and no
   `as_of` — check a dual-licensed domain singly for its other links. An entry we
   can't look up comes back with `error` and **no `verdict`** — report it as not
   checked, never as not found — and doesn't fail the batch.
4. `checked_jurisdictions` and `_meta.stale_jurisdictions` are returned once
   at the top — use them to scope every "not found" verdict. See
   [Batch domain check](https://igregulator.io/docs/batch/).

### Retrospective transaction check (as_of)

"Was this merchant licensed at the time of the transaction?":

1. `GET /v1/check?domain=X&as_of=2026-03-01` (or `/v1/licenses/{id}?as_of=`).
   `as_of` is a `YYYY-MM-DD` date (end of that day, UTC; today's date means now)
   or an ISO-8601 datetime with a UTC offset — anything else, an impossible date
   or a future one is a `400`.
2. `verdict` still describes today; `verdict_detail` leads with the answer for
   your date ("On 2026-03-01 we were not yet tracking … licence … (we first recorded
   it on …), so its status then is unknown. Today: …"). On `/v1/check` the `as_of` object names the
   licence it answered for (`license_id`, `license_number`, `operator`,
   `scope: "licence"`): today's licence for that domain. We keep no record of when
   a domain was linked to a licence (`link_note` says so), so this is that
   licence's status on the date — not proof of who ran the site then.
3. Read the `as_of` object, and **honour `knowledge`**:
   - `observed` → `status_as_of` is the real status then; `established_by`
     shows when it was last confirmed relative to your date.
   - `before_tracking` → the date predates our observation window. Do **not**
     assert a status — tell the user we weren't watching before
     `tracking_since`. This is the difference between a defensible answer and
     a fabricated one.
   - `corrected` → what we recorded for that date was later withdrawn.
     `status_as_of` is `null`; `correction.withdrawn_status` is what we
     **used** to say, never the status. Treat it like `before_tracking`.
   - `no_such_license` → we hold no history for that licence. `null` status.
   - `no_license_resolved` (`/v1/check` only) → the match didn't resolve to
     a specific licence, so there is nothing to time-travel. `null` status.

   Only `observed` carries a status. See
   [Point-in-time lookups](https://igregulator.io/docs/point-in-time/).

## Start integrating

1. Fetch [llms-full.txt](https://igregulator.io/llms-full.txt) for complete docs context in
   a single request.
2. Review the [OpenAPI spec](https://api.igregulator.io/openapi.json).
3. Try endpoints interactively in the
   [playground](https://igregulator.io/docs/playground/).
4. [Create a free account](https://app.igregulator.io/signup) for an
   API key and MCP server access — free for founding members.

---

Questions, integration help, feedback — founder@igregulator.io.
