# Seadvise — crawler policy # # Split by crawler FUNCTION, not by vendor. Measured citation retention after # blocking is highest for pure-training bots (Google-Extended 92.3%, GPTBot # 88.2%) and lowest for the retrieval bot (ChatGPT-User 70.6%). Blocking # training is close to free; blocking retrieval is what costs citations. # # Two things this file cannot do: # - Disallow governs collection, not retention. It cannot remove content # already in model weights or Common Crawl archives. # - Blocking Google-Extended does NOT opt out of AI Overviews or AI Mode, # which draw from the live Search index. data-nosnippet around specific # banded values is the only surgical control. # # ── READ THIS BEFORE ADDING A GROUP ───────────────────────────────────────── # # A crawler obeys exactly ONE group: the most specific User-agent match. It does # NOT fall back to `User-agent: *` for rules its own group omits. # # That cost us every restriction in this file. Googlebot, bingbot and all eight # retrieval crawlers each had a group containing only `Allow: /`, so the # eighteen Disallow rules under `*` applied to nobody except unnamed bots — # including the token routes this file says "must never be indexed", and the # authenticated application surfaces. Fixed 18 Sep 2026 by giving the named # crawlers ONE shared group that carries the restrictions as well as the # permission. # # So: if you add a named User-agent below, it needs the full rule set, not just # the line you came to write. # === Search indexing and AI retrieval: allow the public surface. ============ # # This is the SEO, and retrieval is what produces citations — so these crawlers # get everything public. They are listed together precisely so the restrictions # below cannot drift apart from the permission above them. User-agent: Googlebot User-agent: bingbot User-agent: OAI-SearchBot User-agent: ChatGPT-User User-agent: Claude-SearchBot User-agent: Claude-User User-agent: PerplexityBot User-agent: Perplexity-User User-agent: Applebot User-agent: DuckAssistBot User-agent: * # Public marketing and pricing pages. These must precede the /shipping/ and # /cargo-owner/ disallows below, which would otherwise block them. Allow: /ports/pricing Allow: /shipping/pricing # NOTE: there are deliberately no `Allow:` lines for /ports, /hamnar/ or the # four sector pages. Everything not disallowed is already allowed, and adding # them is not merely redundant — it is actively harmful. # # `Allow: /ports` and `Disallow: /*?q=` are both six characters. Google resolves # an equal-length conflict in favour of Allow, so that one line would have # unblocked the entire facet parameter space the rule below exists to close. # Written down because it was added, caught by the group-semantics test, and # removed again on 18 Sep 2026. # Authenticated application surfaces. # # The trailing slashes are load-bearing. `/shipping/` and `/cargo-owner/` block # the authenticated portals while leaving the marketing pages at `/shipping` # and `/cargo-owner` crawlable — those are server-rendered and worth indexing. Disallow: /admin Disallow: /login Disallow: /port/ Disallow: /shipping/ Disallow: /cargo-owner/ # Registration and token flows — no indexing value, and the token routes must # never be indexed. Disallow: /register Disallow: /cargo-owner-register Disallow: /shipping-register Disallow: /confirm/ Disallow: /restore/ Disallow: /sentry # Query surface — the dataset, not the entities. Blocking the facet parameter # space is also good SEO: it prevents index bloat from combinatorial URLs. # # The port directory's own filtered views additionally carry # `` and a canonical back to # /ports. Both exist on purpose: the disallow stops the crawl, the canonical # handles anything already indexed or reached from an external link, which a # disallow cannot. Disallow: /search Disallow: /export Disallow: /api/ Disallow: /*?q= Disallow: /*?filter= Disallow: /*?sort= Disallow: /*&sort= # === AI training: disallow. Measured citation cost is near zero. ============ # TDM rights reserved under Art. 4(3) Directive (EU) 2019/790 and # 15 a § lag (1960:729) om upphovsrätt. Policy: https://seadvi.se/tdm-policy # # These groups are deliberately a bare `Disallow: /` — there is nothing to # permit, so the shared rule set above does not apply and must not be copied # here. User-agent: GPTBot Disallow: / User-agent: ClaudeBot Disallow: / User-agent: CCBot Disallow: / User-agent: Google-Extended Disallow: / User-agent: Applebot-Extended Disallow: / User-agent: Bytespider Disallow: / User-agent: Meta-ExternalAgent Disallow: / User-agent: Amazonbot Disallow: / Sitemap: https://seadvi.se/sitemap.xml