Data

Loudest reads 428 news outlets in 99 countries every ten minutes, groups the articles into stories, and ranks each story by how many countries and outlets are covering it. Since 22 September 2026 each story is also judged hard news or not, and the map shows only those judged within the rule; the files keep both, with the judgement. Every completed day is published as a snapshot you can download.

What a day contains

One file per UTC day: the stories that spread furthest that day, ranked by the highest point each reached that day, each with the outlets and countries covering it. Days archived before 4 October 2026 are ranked as they stood at midnight UTC. Stories are clusters of articles, not articles. We link to the lead article at its publisher and never reproduce article text, so the dataset is coverage measurement (who covered what, and how widely) rather than a news corpus. The outlets read are listed on the sources page, with each feed's country, language, ownership and class; every story in a day file has its own page at /story/ followed by its id.

How often it updates, and how far back it goes

The reader runs every ten minutes, all day. A day's snapshot is captured shortly after that day ends at midnight UTC, and its JSON and CSV files are written once, at the same moment. From then on the day is fixed: every request for a file returns the bytes written then, and the database refuses any change to a captured day or its files. A day's page may change its design; its files may not. A mistake found later is never corrected in place: it is noted, with its date, on the day's page. The one exception is removal on legal grounds, such as a court order or a valid takedown notice, and every such removal is logged with its reason.

Until 25 September 2026 the files were rebuilt on each request. The files of every day before it were written once on that day, as the code rendered them then, so a copy downloaded earlier may differ. 14 to 18 September 2026 were also edited after capture, before the lock; each carries an erratum, and the changelog records the edits. Days marked partial were not fully covered and should be treated as a sample.

21 days, from 14 September 2026 to 4 October 2026. A new day is added a few minutes after it ends, UTC.

Downloads

Replace the date in either URL with any day in the archive:

https://www.loudest.news/archive/YYYY-MM-DD.json
https://www.loudest.news/archive/YYYY-MM-DD.csv

For example, 2026-10-04.json or 2026-10-04.csv. Both are served straight from the CDN and cached for a year. The CSV is UTF-8 with a header row and CRLF endings, one row per story; the day's own fields repeat on every row, so files for many days concatenate into a single table without losing which day each row came from.

Checking a file

Each file's SHA-256 is on its day page and below, so you can check that a copy you hold is the one published: shasum -a 256 loudest-2026-10-04.csv on macOS, sha256sum on Linux. A citation that gives the SHA-256 of the file used points at exactly those bytes.

SHA-256 of every file, 21 days
DayFileSHA-256
2026-10-04CSVcec35199ec286a729972dce7e70c49f626688dc776f3cb7121af7b7e32067808
2026-10-04JSON6c3265a17e83bc7338f3c9c394706257a162b9064797222a87e66ba5d111a5e0
2026-10-03CSV406bd509d60c6a9c6a089a7044bf8a17e48db474416cfd6841dab77e99c3da94
2026-10-03JSON092c70c60290cdc4fd4673f64158649939b263893cf90789d62773eadda270d7
2026-10-02CSV7850b148c38f8712e97dec0e70133bdaea4abfebee3c780a3af6899e78a89874
2026-10-02JSONe695cfb33313543613bf1461321f6e994671dcab78be3d567b7e971367dc017e
2026-10-01CSVc2caab11de651c792ca38c33d06920fbd0be3220dd42a9a1ce730d63bd5fa898
2026-10-01JSONb1aebec3c5c4358f5482f45b2f72bdcc818394252ac33162b1db4f0dffa46082
2026-09-30CSVa1ae8986c4dde3edbc5eef20ea4a5e50e67c2855ce65629a9304cc4e61335568
2026-09-30JSONe393c6b9288cfc704d8fd7a5863e8540ef60ea48c2a6c79fe794054dd9f4d69c
2026-09-29CSVe42ac9f5fdc59f676caebbd8236c54d8b05416fda6016e137cc48da03b9f7e00
2026-09-29JSON2642b39ab0594169662f3711d0c4b87498838e882e5c935b465625a57ce5f281
2026-09-28CSV69979e2151a393692955c6d3c1645c0b9c78c6a8963e576a39a0ecf13eeaeff6
2026-09-28JSONd18f3de12869679b760817526bd849b81b7169c2af9d2d2f89f1f372bef5e22e
2026-09-27CSV41f45763e0abe75e9bcc17aed6aab47693619956307510a7745dcce595bff5ef
2026-09-27JSON20191837838083e3bb7836b93816553bd499f1271cb4412d049c21d0851e3ba5
2026-09-26CSV3d95b7f445df6abea1293bbe8e5331da4a31386ad0728761e3ad8584518592ec
2026-09-26JSONc1c68d51dad06b1d48730ed0fb48575369ffa86822120c1874bff41d6ab19b8d
2026-09-25CSVbd2f6eaf1418d4285581c61db06075b64d857f939955e51413d477faefaac979
2026-09-25JSONad3e0376dca6c123d9690b5224d7e53925012461ce0db64e97d74ae09c09ed96
2026-09-24CSV19b830e68b4239c1af101088258eb99e34963a4c46a7c11c6f449208f9338403
2026-09-24JSONeea21a88b2cf2623ee80c69a1fb48a5717b8e3aa25c59b09f51c15e71f3d42c5
2026-09-23CSV6fd7fd7e3622c94932ea4a69fa614d476fe4c96f583ad4127a5d4a2fcb0ef600
2026-09-23JSON0c3913f493544b772c1b5e3622c745ed4fd1542b951460a6a477c33ad63b79e1
2026-09-22CSVcc0a81ea9c0e2fd18285057a05cf55901de0f0bd4f4b6d8021feaa95273445b0
2026-09-22JSONc6c0bacc1cc503e0e6a3fec2bd0a55b11c8440c1145e9b4c3c0937c62326cea1
2026-09-21CSVc22998abd5371def93c069a33bc55651b5272ecdcdbe3af9bb9e93d895c77f50
2026-09-21JSONf1b086e424c96bcee565113286d5377b373d856bf9b9aea718ce926614ebc2ef
2026-09-20CSVbc8ed1bdc6ff6609522e14095ef481a72054e21eb87559d3dea4331863a0fbd6
2026-09-20JSON372e5e98746b31143577330c8487dd76df09b365a68044a9c70538d89c96f84c
2026-09-19CSVbf9de9f870fd5b5c49281cb607455b284a86ea68339fc6ff9288eeb1f8b38cec
2026-09-19JSON9ea81d47dcf62de9a592ceb289ebb7a8d4bdeed1c6f13721a1167d12bf0d3861
2026-09-18CSVa2e237e5d7d3971f7e4b82b2ae534d44c88cc78424c9cd6068c5fb9f581b06bc
2026-09-18JSON82dca4553cb796f7577382982dc3369f8082b3f71ff4a2ca77e0a4d660219f1d
2026-09-17CSV1e55e09709be638875191236106e9c32b607b95666eb343feeda3670f38b0970
2026-09-17JSON9909f71d7d2f0c0dfbbb2d269c9ef623b965bffc6fab83b18a40b149f207b2ac
2026-09-16CSV2602397ecd995ff40317adeee28a951aab7bea2d7de93f5ecb282865199a222a
2026-09-16JSONd44423a26db406d1ade0e9fe595cfb93bae6ddff12e2afe62c3ec8ef2a6df016
2026-09-15CSV82908adb8a12339abc06bebe8df4d6ea748acd67b171e3e9fda9fb5e7b2ae716
2026-09-15JSON85428943a3642087bc0edccbec5134b4a8ad0b72fc77b0256c47201f46570c8c
2026-09-14CSVe8ae4a7f21a4988c4dcc5f80803a692cd6571517ffc3f7911f032875fb3e1ff8
2026-09-14JSON02f2ab82e333d53276afb4ff74eb2802ba734d6f86225e05eb116067597ab25b

Day fields

In JSON these sit at the top level, beside a stories array. In CSV they are the first columns of every row.

FieldTypeMeaning
datestringThe UTC day the snapshot covers, YYYY-MM-DD.
partialbooleanTrue when the day was not fully covered: the reader missed runs, started late or stopped early. Treat those days as a sample, not a census.
storyCountintegerStories in this file.
candidateCountinteger | nullStories covered by more than one outlet that day, before the cap on how many are kept. Null only on a day captured before this was recorded; no published day has one.
outletsReadinteger | nullNews outlets the reader read that day: the feeds in its list, less those known to be failing. The same figure the site shows. Null only on a day captured before this was recorded; no published day has one.
countriesReadinteger | nullCountries actually heard from that day: the countries of the outlets that returned at least one article. Outlets with no country of their own are not counted. Null only on a day captured before this was recorded; no published day has one.
languagesReadinteger | nullDistinct languages the day’s stories appeared in. Null only on a day captured before this was recorded; no published day has one.
failedFeedCountinteger | nullOutlets whose feed failed at least once that day, so their coverage is missing or thin. Null only on a day captured before this was recorded; no published day has one.
unreconstructedStoryCountinteger | nullStories in this file whose coverage lists could not be rebuilt, so their countries, languages, outlets, regionCount and articleCount are empty. On a day captured within minutes of its end their counts and score are unaffected; on a day captured later, see contaminatedStoryCount. Null only on a day captured before this was recorded; no published day has one.
contaminatedStoryCountinteger | nullStories in this file whose counts, score, rank, times and lead article include another story’s articles: the day was captured after that story was folded into this one, so the figures are the merged story’s, not what this story alone had at the day’s end. The originals cannot be recovered and are not estimated. Zero on every day captured within minutes of its end. Null only on a day captured before this was recorded; no published day has one.
omittedStoryCountinteger | nullStories folded into another story between the day’s end and its capture, so they are absent from this file and their articles are counted under the story they joined. Zero on every day captured within minutes of its end. Null only on a day captured before this was recorded; no published day has one.
storiesLeftOffinteger | nullStories in this file judged outside the hard-news rule and so left off the map that day: the rows whose hardNews is false. They keep their rank, counts and score here, ranked among all the day’s stories; the map showed the rows whose hardNews is true. Null on days measured under a methodology before 3.0, when every story was shown.
articlesSetAsideinteger | nullArticles set aside before processing that day because their address or feed category clearly marked them as sport, entertainment, arts or lifestyle: never embedded, clustered or counted, so they are in no story in this file. Null on days measured under a methodology before 3.1, when nothing was set aside.
generatedAtstringWhen the snapshot was captured, ISO 8601 UTC. Shortly after the day ended, except for 14, 15 and 16 September 2026, captured together on the morning of the 17th.
methodologyVersionstringThe methodology version in force at the end of the day, major.minor, as listed at /methodology/changelog. A major changes what a figure means; compare days across one with care.
methodologyVersionsstring[]Every methodology version in force during the day, in order. One entry on most days; more on a day a version went live, with the time in the changelog.
In CSV: Joined with a vertical bar, e.g. 1.4|2.0.
methodologyTransitionbooleanTrue when a major version changed during the day, so the day was measured partly under one system and partly under another. The changelog says what changed and at what time. Such a day is neither the day before nor the day after.

Day lists

Also at the top level in JSON. These are the only fields the CSV leaves out: each would repeat on every one of the day's rows, and failedFeeds alone would add hundreds of kilobytes to the file. The day fields above carry the length of each, so a CSV reader still knows what is there.

FieldTypeMeaning
outletNamesstring[] | nullEvery outlet covering the day’s stories, sorted. Each story’s outlets are positions in this list. Null only on a day captured before this was recorded; no published day has one.
languagesstring[] | nullEvery language the day’s stories appeared in, sorted, ISO 639-1. Null only on a day captured before this was recorded; no published day has one.
failedFeedsstring[] | nullOutlets whose feed failed at least once that day, sorted. A feed that failed all day contributed nothing, so the day’s counts understate its country and language. Null only on a day captured before this was recorded; no published day has one.
unreconstructedStoryIdsstring[] | nullThe ids of the stories counted by unreconstructedStoryCount. Null only on a day captured before this was recorded; no published day has one.
contaminatedStoryIdsstring[] | nullThe ids of the stories counted by contaminatedStoryCount. Null only on a day captured before this was recorded; no published day has one.
omittedStoriesobject[] | nullOne entry per story counted by omittedStoryCount: id, title, mergedInto (the id of the story it was folded into, usually one in this file), mergedAt, mergeId, and outletCount and countryCount as they stood at the merge, after the day’s end. Whether it had ranked that day, and where, is not known. Null only on a day captured before this was recorded; no published day has one.

Story fields

The site’s own terms, first seen and first published among them, are defined in the glossary.

FieldTypeMeaning
rankintegerPosition in the day, 1 being the most widely covered. Ranked by the highest score the story reached that day. On days measured under a methodology before 4.1, ranked by score at the day’s end.
idstringStable identifier for the story cluster. The same cluster keeps its id across days.
titlestringA neutral headline written for the cluster, not copied from any one outlet.
urlstringThe lead article, at its publisher. We do not host article text.
sourcestringThe outlet that published the lead article.
regionstringWhere the story is about: Africa, Americas, Asia-Pacific, Europe, Middle East, or World when it is not one place. World is the value in every file and API response; the site labels it International.
countrystring | nullWhere the story is about, ISO 3166-1 alpha-2. Null when it is not about a single country.
topicstring | nullTopic identifier, e.g. politics or conflict. Null when none was assigned.
hardNewsboolean | nullThe judgement as applied: true when the story’s subject was judged within the hard-news rule and it could appear on the map, false when it was judged outside and left off. Made from the same headlines as the title, by the same model call, and weighed against the judgements of similar recent stories. Null for a story never judged: every story on days measured under a methodology before 3.0, and a story that was never covered widely enough to be judged.
judgementSureboolean | nullWhether the judgement was sure of itself. An unsure judgement on a widely covered story is shown rather than hidden, so hardNews true with judgementSure false marks a story shown on that rule. Null whenever hardNews is null.
outletCountintegerDistinct outlets covering the story, general and trade press together.
countryCountintegerDistinct countries those outlets publish from.
generalOutletCountintegerDistinct general-press outlets covering it, excluding trade and specialist titles.
generalCountryCountintegerDistinct countries of those general-press outlets. This is the main term in the score.
outletCount24hinteger | nullDistinct outlets that covered the story in the last 24 hours of the day, general and trade press together: the day’s coverage, where outletCount is the cluster’s whole life. Null on days before 21 September 2026, when it was not recorded.
countryCount24hinteger | nullDistinct countries those outlets publish from, over the same 24 hours. Null on days before 21 September 2026.
scorenumberThe highest coverage score the story reached that day: one vote per general-press country, plus a smaller term for outlets, each fading from when it joined. Worked out when the day is archived, from when each country and outlet joined the story, so a story merged with another during the day is scored as the two together and can stand higher than either did on the map. On days measured under a methodology before 4.1, the score at the day’s end, faded by how long ago the story was last seen.
peakAtstring | nullWhen the story reached that score, ISO 8601 UTC: the start of the reader’s run that measured it, the earliest if more than one did. Null for a story no run of the day scored above zero. Null on days measured under a methodology before 4.1, when it was not recorded.
firstSeenAtstringMisnamed: the same value as firstPublishedAt, kept for compatibility with files already in use. On the site, “first seen” means the first fetch, which is firstFetchedAt; this field is the earliest publisher date. Read firstPublishedAt instead.
firstFetchedAtstring | nullWhen the reader first fetched an article in the cluster, ISO 8601 UTC: the first time it saw the story, and the age the site shows. Fetched, not published, so it trails firstSeenAt by the reader’s own delay, ten minutes between runs. Null only on a day captured before this was recorded, for a story that could not be rebuilt afterwards.
lastSeenAtstringWhen the reader last saw a new article in the cluster, ISO 8601 UTC. The date the publisher gave, or the fetch time when it gave none.
countriesstring[] | nullThe countries the outlets covering it publish from, ISO 3166-1 alpha-2, sorted. Shorter than countryCount when an outlet has no country of its own; AllAfrica and UN News are pan-national, and countryCount counts them together as one unknown.
In CSV: The codes joined with a vertical bar, e.g. AU|FR|US.
languagesstring[] | nullThe languages it was covered in, ISO 639-1, sorted. The language of the outlet’s feed, not of the story.
In CSV: The tags joined with a vertical bar, e.g. de|en|es.
regionCountinteger | nullHow many of the six world regions the outlets covering it publish from. The spread the region field does not give: region says where the story is about, this says how far it travelled.
articleCountinteger | nullArticles in the cluster. Higher than outletCount when outlets filed more than once, so the two together separate wide coverage from repeated coverage.
outletsinteger[] | nullThe outlets covering it, as positions in the day’s outletNames list. Indices rather than names: a day names a couple of hundred outlets and cites them thousands of times, so spelling each one out would add about a fifth to the file.
In CSV: Resolved to names and joined with a vertical bar, e.g. BBC News|CBS News, since the CSV has no outletNames list to point into.
firstPublishedAtstringThe earliest date a publisher gave for any article in the cluster, ISO 8601 UTC. A publisher’s claim, not an observation: a feed can carry an item days after it was written, and this precedes firstFetchedAt by more than a day for about 5% of stories. The exception is 14 September 2026, the first day, where it is a fifth: the first run read every feed’s backlog of up to 48 hours at once, so that day’s first fetches all start at 17:15 UTC. When the earliest article carries no date this is its fetch time, equal to firstFetchedAt. May precede the day itself for a story that ran on. Added 21 September 2026 as the correct name for firstSeenAt, which carries the same value.
continuesstring | nullThe id of the story this one continues. A story accepts articles for 48 hours after its first; an article that arrives later and resembles it starts a new story that records this one as its predecessor, so a narrative that runs on is a chain of stories rather than one growing cluster. Null when the story resembled nothing recent when it began. Null also on days before 21 September 2026, when it was not recorded, so those days cannot be chained.

Lists in the CSV

CSV has no way to nest a list inside a cell, so each of the story lists is flattened into a single field, its values joined with a vertical bar: |. No outlet name, language tag or country code in the dataset contains one, so splitting on it is lossless. In Python: row['countries'].split('|').

outlets is the exception worth knowing about. In JSON it holds positions in the day's outletNames list, because writing 215 outlet names out across 6,000-odd mentions would be a fifth of the file. The CSV has no list to point into, so it resolves them and writes the names in full.

An empty list field means the story's coverage lists could not be rebuilt, not that nothing covered it; see unreconstructedStoryCount. Every story that was reconstructed has at least one outlet.

Why some stories have no lists

The lists were added on 18 September 2026 and are captured with each day from then on. The days before it were filled in from the articles they were built from, which survive a week past a story's last article. Two kinds of story could not be filled in: one that was folded into another story after the day ended, whose row no longer exists, and the story it was folded into, whose articles now include the other's and so no longer say what it alone covered that day. Rather than publish a list that disagrees with the outletCount beside it, both are left empty and counted in unreconstructedStoryCount. On a day captured within minutes of its end, their counts, score and ranking are the originals and are unaffected. On a day captured later they may not be, which is the next section.

Days captured late

A snapshot is computed from the articles as they are grouped into stories at the moment of capture. When two stories about one event are merged, the survivor takes the other's articles and the other's row is deleted. A day captured within minutes of its end is taken before any such merge can touch it, and since 19 September 2026 the capture runs before the merge pass in the first run after midnight, so on those days contaminatedStoryCount and omittedStoryCount are zero.

14, 15 and 16 September 2026 were captured together on the morning of the 17th, after six merges, and show both effects. Nine stories across the three days carry the merged story's figures rather than their own at the day's end: outletCount, countryCount, the general-press counts, score, rank, the three times and the lead article, and a headline written after the merge. They are listed in contaminatedStoryIds. The stories folded away are absent from those files, and are listed in omittedStories with the id of the story each joined and the figures the merge log recorded for it, which are as at the merge, not at the day's end. The originals cannot be recovered, because nothing records which articles were whose, so they are flagged rather than estimated. A contaminated story is also counted as unreconstructed, since its lists could not be rebuilt for the same reason.

Licence

The daily data files, the JSON and CSV file for each day, and the JSON record of each story, are published under Creative Commons Attribution 4.0 International (CC BY 4.0). You may use them for anything, including commercially, as long as you credit them. Attribute as: Data from loudest.news, CC BY 4.0., with a link to loudest.news where the medium allows.

The licence covers the measurements in those files: the clusters, counts, scores, rankings and headlines written for each cluster. It does not cover the articles we link to, which belong to the outlets that published them, and it does not cover the rest of the site, which is reserved as the terms describe.

Ready-made citations, plain text and BibTeX, for the dataset, a day, a story and the methodology are on the cite page.