Data
Loudest reads 428 news outlets in 99 countries every ten minutes, groups the articles into stories, and ranks each story by how many countries and outlets are covering it. Since 22 September 2026 each story is also judged hard news or not, and the map shows only those judged within the rule; the files keep both, with the judgement. Every completed day is published as a snapshot you can download.
What a day contains
One file per UTC day: the stories that spread furthest that day, ranked by the highest point each reached that day, each with the outlets and countries covering it. Days archived before 4 October 2026 are ranked as they stood at midnight UTC. Stories are clusters of articles, not articles. We link to the lead article at its publisher and never reproduce article text, so the dataset is coverage measurement (who covered what, and how widely) rather than a news corpus. The outlets read are listed on the sources page, with each feed's country, language, ownership and class; every story in a day file has its own page at /story/ followed by its id.
How often it updates, and how far back it goes
The reader runs every ten minutes, all day. A day's snapshot is captured shortly after that day ends at midnight UTC, and its JSON and CSV files are written once, at the same moment. From then on the day is fixed: every request for a file returns the bytes written then, and the database refuses any change to a captured day or its files. A day's page may change its design; its files may not. A mistake found later is never corrected in place: it is noted, with its date, on the day's page. The one exception is removal on legal grounds, such as a court order or a valid takedown notice, and every such removal is logged with its reason.
Until 25 September 2026 the files were rebuilt on each request. The files of every day before it were written once on that day, as the code rendered them then, so a copy downloaded earlier may differ. 14 to 18 September 2026 were also edited after capture, before the lock; each carries an erratum, and the changelog records the edits. Days marked partial were not fully covered and should be treated as a sample.
21 days, from 14 September 2026 to 4 October 2026. A new day is added a few minutes after it ends, UTC.
Downloads
Replace the date in either URL with any day in the archive:
https://www.loudest.news/archive/YYYY-MM-DD.jsonhttps://www.loudest.news/archive/YYYY-MM-DD.csv
For example, 2026-10-04.json or 2026-10-04.csv. Both are served straight from the CDN and cached for a year. The CSV is UTF-8 with a header row and CRLF endings, one row per story; the day's own fields repeat on every row, so files for many days concatenate into a single table without losing which day each row came from.
Checking a file
Each file's SHA-256 is on its day page and below, so you can check that a copy you hold is the one published: shasum -a 256 loudest-2026-10-04.csv on macOS, sha256sum on Linux. A citation that gives the SHA-256 of the file used points at exactly those bytes.
SHA-256 of every file, 21 days
| Day | File | SHA-256 |
|---|---|---|
| 2026-10-04 | CSV | cec35199ec286a729972dce7e70c49f626688dc776f3cb7121af7b7e32067808 |
| 2026-10-04 | JSON | 6c3265a17e83bc7338f3c9c394706257a162b9064797222a87e66ba5d111a5e0 |
| 2026-10-03 | CSV | 406bd509d60c6a9c6a089a7044bf8a17e48db474416cfd6841dab77e99c3da94 |
| 2026-10-03 | JSON | 092c70c60290cdc4fd4673f64158649939b263893cf90789d62773eadda270d7 |
| 2026-10-02 | CSV | 7850b148c38f8712e97dec0e70133bdaea4abfebee3c780a3af6899e78a89874 |
| 2026-10-02 | JSON | e695cfb33313543613bf1461321f6e994671dcab78be3d567b7e971367dc017e |
| 2026-10-01 | CSV | c2caab11de651c792ca38c33d06920fbd0be3220dd42a9a1ce730d63bd5fa898 |
| 2026-10-01 | JSON | b1aebec3c5c4358f5482f45b2f72bdcc818394252ac33162b1db4f0dffa46082 |
| 2026-09-30 | CSV | a1ae8986c4dde3edbc5eef20ea4a5e50e67c2855ce65629a9304cc4e61335568 |
| 2026-09-30 | JSON | e393c6b9288cfc704d8fd7a5863e8540ef60ea48c2a6c79fe794054dd9f4d69c |
| 2026-09-29 | CSV | e42ac9f5fdc59f676caebbd8236c54d8b05416fda6016e137cc48da03b9f7e00 |
| 2026-09-29 | JSON | 2642b39ab0594169662f3711d0c4b87498838e882e5c935b465625a57ce5f281 |
| 2026-09-28 | CSV | 69979e2151a393692955c6d3c1645c0b9c78c6a8963e576a39a0ecf13eeaeff6 |
| 2026-09-28 | JSON | d18f3de12869679b760817526bd849b81b7169c2af9d2d2f89f1f372bef5e22e |
| 2026-09-27 | CSV | 41f45763e0abe75e9bcc17aed6aab47693619956307510a7745dcce595bff5ef |
| 2026-09-27 | JSON | 20191837838083e3bb7836b93816553bd499f1271cb4412d049c21d0851e3ba5 |
| 2026-09-26 | CSV | 3d95b7f445df6abea1293bbe8e5331da4a31386ad0728761e3ad8584518592ec |
| 2026-09-26 | JSON | c1c68d51dad06b1d48730ed0fb48575369ffa86822120c1874bff41d6ab19b8d |
| 2026-09-25 | CSV | bd2f6eaf1418d4285581c61db06075b64d857f939955e51413d477faefaac979 |
| 2026-09-25 | JSON | ad3e0376dca6c123d9690b5224d7e53925012461ce0db64e97d74ae09c09ed96 |
| 2026-09-24 | CSV | 19b830e68b4239c1af101088258eb99e34963a4c46a7c11c6f449208f9338403 |
| 2026-09-24 | JSON | eea21a88b2cf2623ee80c69a1fb48a5717b8e3aa25c59b09f51c15e71f3d42c5 |
| 2026-09-23 | CSV | 6fd7fd7e3622c94932ea4a69fa614d476fe4c96f583ad4127a5d4a2fcb0ef600 |
| 2026-09-23 | JSON | 0c3913f493544b772c1b5e3622c745ed4fd1542b951460a6a477c33ad63b79e1 |
| 2026-09-22 | CSV | cc0a81ea9c0e2fd18285057a05cf55901de0f0bd4f4b6d8021feaa95273445b0 |
| 2026-09-22 | JSON | c6c0bacc1cc503e0e6a3fec2bd0a55b11c8440c1145e9b4c3c0937c62326cea1 |
| 2026-09-21 | CSV | c22998abd5371def93c069a33bc55651b5272ecdcdbe3af9bb9e93d895c77f50 |
| 2026-09-21 | JSON | f1b086e424c96bcee565113286d5377b373d856bf9b9aea718ce926614ebc2ef |
| 2026-09-20 | CSV | bc8ed1bdc6ff6609522e14095ef481a72054e21eb87559d3dea4331863a0fbd6 |
| 2026-09-20 | JSON | 372e5e98746b31143577330c8487dd76df09b365a68044a9c70538d89c96f84c |
| 2026-09-19 | CSV | bf9de9f870fd5b5c49281cb607455b284a86ea68339fc6ff9288eeb1f8b38cec |
| 2026-09-19 | JSON | 9ea81d47dcf62de9a592ceb289ebb7a8d4bdeed1c6f13721a1167d12bf0d3861 |
| 2026-09-18 | CSV | a2e237e5d7d3971f7e4b82b2ae534d44c88cc78424c9cd6068c5fb9f581b06bc |
| 2026-09-18 | JSON | 82dca4553cb796f7577382982dc3369f8082b3f71ff4a2ca77e0a4d660219f1d |
| 2026-09-17 | CSV | 1e55e09709be638875191236106e9c32b607b95666eb343feeda3670f38b0970 |
| 2026-09-17 | JSON | 9909f71d7d2f0c0dfbbb2d269c9ef623b965bffc6fab83b18a40b149f207b2ac |
| 2026-09-16 | CSV | 2602397ecd995ff40317adeee28a951aab7bea2d7de93f5ecb282865199a222a |
| 2026-09-16 | JSON | d44423a26db406d1ade0e9fe595cfb93bae6ddff12e2afe62c3ec8ef2a6df016 |
| 2026-09-15 | CSV | 82908adb8a12339abc06bebe8df4d6ea748acd67b171e3e9fda9fb5e7b2ae716 |
| 2026-09-15 | JSON | 85428943a3642087bc0edccbec5134b4a8ad0b72fc77b0256c47201f46570c8c |
| 2026-09-14 | CSV | e8ae4a7f21a4988c4dcc5f80803a692cd6571517ffc3f7911f032875fb3e1ff8 |
| 2026-09-14 | JSON | 02f2ab82e333d53276afb4ff74eb2802ba734d6f86225e05eb116067597ab25b |
Day fields
In JSON these sit at the top level, beside a stories array. In CSV they are the first columns of every row.
| Field | Type | Meaning |
|---|---|---|
date | string | The UTC day the snapshot covers, YYYY-MM-DD. |
partial | boolean | True when the day was not fully covered: the reader missed runs, started late or stopped early. Treat those days as a sample, not a census. |
storyCount | integer | Stories in this file. |
candidateCount | integer | null | Stories covered by more than one outlet that day, before the cap on how many are kept. Null only on a day captured before this was recorded; no published day has one. |
outletsRead | integer | null | News outlets the reader read that day: the feeds in its list, less those known to be failing. The same figure the site shows. Null only on a day captured before this was recorded; no published day has one. |
countriesRead | integer | null | Countries actually heard from that day: the countries of the outlets that returned at least one article. Outlets with no country of their own are not counted. Null only on a day captured before this was recorded; no published day has one. |
languagesRead | integer | null | Distinct languages the day’s stories appeared in. Null only on a day captured before this was recorded; no published day has one. |
failedFeedCount | integer | null | Outlets whose feed failed at least once that day, so their coverage is missing or thin. Null only on a day captured before this was recorded; no published day has one. |
unreconstructedStoryCount | integer | null | Stories in this file whose coverage lists could not be rebuilt, so their countries, languages, outlets, regionCount and articleCount are empty. On a day captured within minutes of its end their counts and score are unaffected; on a day captured later, see contaminatedStoryCount. Null only on a day captured before this was recorded; no published day has one. |
contaminatedStoryCount | integer | null | Stories in this file whose counts, score, rank, times and lead article include another story’s articles: the day was captured after that story was folded into this one, so the figures are the merged story’s, not what this story alone had at the day’s end. The originals cannot be recovered and are not estimated. Zero on every day captured within minutes of its end. Null only on a day captured before this was recorded; no published day has one. |
omittedStoryCount | integer | null | Stories folded into another story between the day’s end and its capture, so they are absent from this file and their articles are counted under the story they joined. Zero on every day captured within minutes of its end. Null only on a day captured before this was recorded; no published day has one. |
storiesLeftOff | integer | null | Stories in this file judged outside the hard-news rule and so left off the map that day: the rows whose hardNews is false. They keep their rank, counts and score here, ranked among all the day’s stories; the map showed the rows whose hardNews is true. Null on days measured under a methodology before 3.0, when every story was shown. |
articlesSetAside | integer | null | Articles set aside before processing that day because their address or feed category clearly marked them as sport, entertainment, arts or lifestyle: never embedded, clustered or counted, so they are in no story in this file. Null on days measured under a methodology before 3.1, when nothing was set aside. |
generatedAt | string | When the snapshot was captured, ISO 8601 UTC. Shortly after the day ended, except for 14, 15 and 16 September 2026, captured together on the morning of the 17th. |
methodologyVersion | string | The methodology version in force at the end of the day, major.minor, as listed at /methodology/changelog. A major changes what a figure means; compare days across one with care. |
methodologyVersions | string[] | Every methodology version in force during the day, in order. One entry on most days; more on a day a version went live, with the time in the changelog. In CSV: Joined with a vertical bar, e.g. 1.4|2.0. |
methodologyTransition | boolean | True when a major version changed during the day, so the day was measured partly under one system and partly under another. The changelog says what changed and at what time. Such a day is neither the day before nor the day after. |
Day lists
Also at the top level in JSON. These are the only fields the CSV leaves out: each would repeat on every one of the day's rows, and failedFeeds alone would add hundreds of kilobytes to the file. The day fields above carry the length of each, so a CSV reader still knows what is there.
| Field | Type | Meaning |
|---|---|---|
outletNames | string[] | null | Every outlet covering the day’s stories, sorted. Each story’s outlets are positions in this list. Null only on a day captured before this was recorded; no published day has one. |
languages | string[] | null | Every language the day’s stories appeared in, sorted, ISO 639-1. Null only on a day captured before this was recorded; no published day has one. |
failedFeeds | string[] | null | Outlets whose feed failed at least once that day, sorted. A feed that failed all day contributed nothing, so the day’s counts understate its country and language. Null only on a day captured before this was recorded; no published day has one. |
unreconstructedStoryIds | string[] | null | The ids of the stories counted by unreconstructedStoryCount. Null only on a day captured before this was recorded; no published day has one. |
contaminatedStoryIds | string[] | null | The ids of the stories counted by contaminatedStoryCount. Null only on a day captured before this was recorded; no published day has one. |
omittedStories | object[] | null | One entry per story counted by omittedStoryCount: id, title, mergedInto (the id of the story it was folded into, usually one in this file), mergedAt, mergeId, and outletCount and countryCount as they stood at the merge, after the day’s end. Whether it had ranked that day, and where, is not known. Null only on a day captured before this was recorded; no published day has one. |
Story fields
The site’s own terms, first seen and first published among them, are defined in the glossary.
| Field | Type | Meaning |
|---|---|---|
rank | integer | Position in the day, 1 being the most widely covered. Ranked by the highest score the story reached that day. On days measured under a methodology before 4.1, ranked by score at the day’s end. |
id | string | Stable identifier for the story cluster. The same cluster keeps its id across days. |
title | string | A neutral headline written for the cluster, not copied from any one outlet. |
url | string | The lead article, at its publisher. We do not host article text. |
source | string | The outlet that published the lead article. |
region | string | Where the story is about: Africa, Americas, Asia-Pacific, Europe, Middle East, or World when it is not one place. World is the value in every file and API response; the site labels it International. |
country | string | null | Where the story is about, ISO 3166-1 alpha-2. Null when it is not about a single country. |
topic | string | null | Topic identifier, e.g. politics or conflict. Null when none was assigned. |
hardNews | boolean | null | The judgement as applied: true when the story’s subject was judged within the hard-news rule and it could appear on the map, false when it was judged outside and left off. Made from the same headlines as the title, by the same model call, and weighed against the judgements of similar recent stories. Null for a story never judged: every story on days measured under a methodology before 3.0, and a story that was never covered widely enough to be judged. |
judgementSure | boolean | null | Whether the judgement was sure of itself. An unsure judgement on a widely covered story is shown rather than hidden, so hardNews true with judgementSure false marks a story shown on that rule. Null whenever hardNews is null. |
outletCount | integer | Distinct outlets covering the story, general and trade press together. |
countryCount | integer | Distinct countries those outlets publish from. |
generalOutletCount | integer | Distinct general-press outlets covering it, excluding trade and specialist titles. |
generalCountryCount | integer | Distinct countries of those general-press outlets. This is the main term in the score. |
outletCount24h | integer | null | Distinct outlets that covered the story in the last 24 hours of the day, general and trade press together: the day’s coverage, where outletCount is the cluster’s whole life. Null on days before 21 September 2026, when it was not recorded. |
countryCount24h | integer | null | Distinct countries those outlets publish from, over the same 24 hours. Null on days before 21 September 2026. |
score | number | The highest coverage score the story reached that day: one vote per general-press country, plus a smaller term for outlets, each fading from when it joined. Worked out when the day is archived, from when each country and outlet joined the story, so a story merged with another during the day is scored as the two together and can stand higher than either did on the map. On days measured under a methodology before 4.1, the score at the day’s end, faded by how long ago the story was last seen. |
peakAt | string | null | When the story reached that score, ISO 8601 UTC: the start of the reader’s run that measured it, the earliest if more than one did. Null for a story no run of the day scored above zero. Null on days measured under a methodology before 4.1, when it was not recorded. |
firstSeenAt | string | Misnamed: the same value as firstPublishedAt, kept for compatibility with files already in use. On the site, “first seen” means the first fetch, which is firstFetchedAt; this field is the earliest publisher date. Read firstPublishedAt instead. |
firstFetchedAt | string | null | When the reader first fetched an article in the cluster, ISO 8601 UTC: the first time it saw the story, and the age the site shows. Fetched, not published, so it trails firstSeenAt by the reader’s own delay, ten minutes between runs. Null only on a day captured before this was recorded, for a story that could not be rebuilt afterwards. |
lastSeenAt | string | When the reader last saw a new article in the cluster, ISO 8601 UTC. The date the publisher gave, or the fetch time when it gave none. |
countries | string[] | null | The countries the outlets covering it publish from, ISO 3166-1 alpha-2, sorted. Shorter than countryCount when an outlet has no country of its own; AllAfrica and UN News are pan-national, and countryCount counts them together as one unknown. In CSV: The codes joined with a vertical bar, e.g. AU|FR|US. |
languages | string[] | null | The languages it was covered in, ISO 639-1, sorted. The language of the outlet’s feed, not of the story. In CSV: The tags joined with a vertical bar, e.g. de|en|es. |
regionCount | integer | null | How many of the six world regions the outlets covering it publish from. The spread the region field does not give: region says where the story is about, this says how far it travelled. |
articleCount | integer | null | Articles in the cluster. Higher than outletCount when outlets filed more than once, so the two together separate wide coverage from repeated coverage. |
outlets | integer[] | null | The outlets covering it, as positions in the day’s outletNames list. Indices rather than names: a day names a couple of hundred outlets and cites them thousands of times, so spelling each one out would add about a fifth to the file. In CSV: Resolved to names and joined with a vertical bar, e.g. BBC News|CBS News, since the CSV has no outletNames list to point into. |
firstPublishedAt | string | The earliest date a publisher gave for any article in the cluster, ISO 8601 UTC. A publisher’s claim, not an observation: a feed can carry an item days after it was written, and this precedes firstFetchedAt by more than a day for about 5% of stories. The exception is 14 September 2026, the first day, where it is a fifth: the first run read every feed’s backlog of up to 48 hours at once, so that day’s first fetches all start at 17:15 UTC. When the earliest article carries no date this is its fetch time, equal to firstFetchedAt. May precede the day itself for a story that ran on. Added 21 September 2026 as the correct name for firstSeenAt, which carries the same value. |
continues | string | null | The id of the story this one continues. A story accepts articles for 48 hours after its first; an article that arrives later and resembles it starts a new story that records this one as its predecessor, so a narrative that runs on is a chain of stories rather than one growing cluster. Null when the story resembled nothing recent when it began. Null also on days before 21 September 2026, when it was not recorded, so those days cannot be chained. |
Lists in the CSV
CSV has no way to nest a list inside a cell, so each of the story lists is flattened into a single field, its values joined with a vertical bar: |. No outlet name, language tag or country code in the dataset contains one, so splitting on it is lossless. In Python: row['countries'].split('|').
outlets is the exception worth knowing about. In JSON it holds positions in the day's outletNames list, because writing 215 outlet names out across 6,000-odd mentions would be a fifth of the file. The CSV has no list to point into, so it resolves them and writes the names in full.
An empty list field means the story's coverage lists could not be rebuilt, not that nothing covered it; see unreconstructedStoryCount. Every story that was reconstructed has at least one outlet.
Why some stories have no lists
The lists were added on 18 September 2026 and are captured with each day from then on. The days before it were filled in from the articles they were built from, which survive a week past a story's last article. Two kinds of story could not be filled in: one that was folded into another story after the day ended, whose row no longer exists, and the story it was folded into, whose articles now include the other's and so no longer say what it alone covered that day. Rather than publish a list that disagrees with the outletCount beside it, both are left empty and counted in unreconstructedStoryCount. On a day captured within minutes of its end, their counts, score and ranking are the originals and are unaffected. On a day captured later they may not be, which is the next section.
Days captured late
A snapshot is computed from the articles as they are grouped into stories at the moment of capture. When two stories about one event are merged, the survivor takes the other's articles and the other's row is deleted. A day captured within minutes of its end is taken before any such merge can touch it, and since 19 September 2026 the capture runs before the merge pass in the first run after midnight, so on those days contaminatedStoryCount and omittedStoryCount are zero.
14, 15 and 16 September 2026 were captured together on the morning of the 17th, after six merges, and show both effects. Nine stories across the three days carry the merged story's figures rather than their own at the day's end: outletCount, countryCount, the general-press counts, score, rank, the three times and the lead article, and a headline written after the merge. They are listed in contaminatedStoryIds. The stories folded away are absent from those files, and are listed in omittedStories with the id of the story each joined and the figures the merge log recorded for it, which are as at the merge, not at the day's end. The originals cannot be recovered, because nothing records which articles were whose, so they are flagged rather than estimated. A contaminated story is also counted as unreconstructed, since its lists could not be rebuilt for the same reason.
Licence
The daily data files, the JSON and CSV file for each day, and the JSON record of each story, are published under Creative Commons Attribution 4.0 International (CC BY 4.0). You may use them for anything, including commercially, as long as you credit them. Attribute as: Data from loudest.news, CC BY 4.0., with a link to loudest.news where the medium allows.
The licence covers the measurements in those files: the clusters, counts, scores, rankings and headlines written for each cluster. It does not cover the articles we link to, which belong to the outlets that published them, and it does not cover the rest of the site, which is reserved as the terms describe.
Ready-made citations, plain text and BibTeX, for the dataset, a day, a story and the methodology are on the cite page.