# LoudestBot

LoudestBot is the feed reader behind [Loudest](https://www.loudest.news), a site that shows
the hard news the world's newsrooms are covering by reading outlets' public RSS and Atom feeds and
sizing each story by how many countries and outlets report it. This page says what the bot
does, how often, how to recognise it, and how to keep it out. Contact: contact@loudest.news.

## What it reads

Only feeds. LoudestBot fetches each outlet's public RSS or Atom feed, listed by hand in the
site's feed list, and nothing else: it does not crawl sites, follow links from feeds, or fetch
article pages. From each feed it keeps the headline, the link to the outlet's own page, the
outlet's name and the publication time. The feed's summary text is held only until the article
has been embedded, in the same run, minutes after the fetch, and the text is then cleared.
Only the headline and a short excerpt of the summary are used to group articles. An item whose address or feed category marks it as
sport, entertainment, arts or lifestyle is dropped at the feed: its address is logged for a
week so it is counted once, and nothing else of it is stored. No article text is republished. Headlines are shown
with a link back to the outlet's own page.

## What it does not do

Loudest does not use what it reads to train AI models. Two outside services see the text:
the headline and a short excerpt of the summary are sent to a service that compares articles by meaning, and the
headlines of a multi-outlet story go to a service that writes one short neutral headline for
it and says whether the story is hard news. Both are used under terms that do not allow training on what is sent. Nothing else is done
with the text.

## How it identifies itself

Every request carries this user agent:

```
LoudestBot/1.0 (+https://www.loudest.news; contact@loudest.news)
```

Every request is also signed under Web Bot Auth (RFC 9421 HTTP message signatures, Ed25519).
The public key is published at:

```
https://www.loudest.news/.well-known/http-message-signatures-directory
```

The requests come from Vercel's serverless platform, so there is no fixed address range; the
signature and the user agent are the identification. The bot never poses as a browser.

## How often and how politely

- Each feed is polled at most once every ten minutes, and less often when the feed asks:
  `Cache-Control: max-age`, RSS `<ttl>` and `sy:updatePeriod` hold the next request back for
  as long as they say, up to 24 hours.
- Requests to one host go one at a time, at least a second apart.
- Where the outlet supports it, requests are conditional: the bot sends `If-None-Match` and
  `If-Modified-Since` and accepts a 304 in place of a body.
- `robots.txt` is read before the first request to a host and again about once a day, and
  obeyed, including `Crawl-delay`. A feed whose path is disallowed for LoudestBot or for `*`
  is not fetched.
- After a failure the bot waits ten minutes before trying again, doubling to six hours while
  the failures continue.
- Each request times out after eight seconds; a slow feed is dropped for that run.

In plain terms: at most one small request per feed every ten minutes. Across about 435 feeds
that is no more than about 63,000 requests a day, and fewer in practice, since some feeds ask
to be read less often.

## How to keep it out

Either of these works:

- A `robots.txt` rule covering the feed's path:

  ```
  User-agent: LoudestBot
  Disallow: /
  ```

  The bot stops fetching the feed at its next read of the file, within about a day, and the
  feed is later moved to the reserve list by hand.

- An email to contact@loudest.news naming the outlet.

A 403 or another refusal is treated differently. The bot backs off and keeps trying at a low
rate, because some blocks are aimed at the shared cloud addresses it runs from rather than at
LoudestBot, and are lifted without the outlet knowing they were there. An outlet that means to
block LoudestBot should use the robots.txt rule or the email address; the feed is then
dropped, not retried. A robots.txt rule or an email covers every feed from that publisher.

## What Loudest does with it

Headlines are grouped into stories across languages and ranked by how widely each is covered.
The result is published live and archived daily as open data under CC BY 4.0, with each story
crediting the outlets that covered it. The [methodology](/methodology) describes the pipeline
in full; the [data page](/data) describes what is published.
