Skip to content
Sources & credit

Who we read, and how we credit them

Huntaegis analyses what publishers report. It does not replace their reporting. This page says what we keep, show, say aloud, translate and post from each story, how our fetcher behaves, and how a publisher can reach us.

What we do with each story

We read 354 feeds from 311 publishers (the full list, grouped by topic, is on the sources page). For each story that enters the pipeline:

Kept
The page text is kept on our server for 30 days so the analysis can read it. It is not shown or served to the public, and it is deleted with the story. One sister learning service of ours can read it to write quiz questions. A translation, if a reader asks for one, is cached for the same period.
Shown
The publisher's headline, a short excerpt (two sentences, at most 400 characters) under “From the original”, our own analysis, and a link to the original, with “Reporting by <publisher>” directly under the headline and the same link under the excerpt. The headline links to the publisher.
Said aloud
Audio is read from our analysis, not from the article. Every narration opens by naming the publisher and ends by saying it was an automated analysis, read by a synthetic voice, of that publisher's reporting. See the voice credits.
Translated
On request, an AI model translates the headline, the excerpt and our analysis. A translated view says so, names the language it came from, and keeps the link to the original.
Posted
Social posts carry the headline, a line of our analysis and a link to our page, and end with “Original reporting: <publisher>” linked to the original story.
Images
See Images below.

How our fetcher behaves

  • It says who it is. Every request carries HuntaegisBot/1.0 (+https://huntaegis.com/about/sources) and never pretends to be a browser.
  • It reads robots.txt before it fetches a feed, an article or an image, follows the Allow/Disallow rules for its own name (or *), and honours Crawl-delay. If the file cannot be reached, it waits instead of assuming yes.
  • It honours noai, nosnippet and noimageai in a page's robots meta tag or X-Robots-Tag header: that text or image is not kept.
  • A refusal is an answer. A 403 or 429 is not retried with other headers, other tools or a headless browser.
  • It waits at least a couple of seconds between requests to the same publisher, and longer if the publisher asks.

What publishers say back (survey of 311 publisher hosts, 2026-10-01): 253 publish a robots.txt; 16 ask for a crawl delay; 3 disallow our reading their feed and 2 disallow their article pages (we do not fetch those); 21 block one or more AI crawlers sitewide.

Publisher terms we enforce

When a publisher asks us not to copy their text or images, the request is written into a configuration file the fetcher reads on every request, so it is enforced rather than remembered. This is that file's current content.

  • apnews.comlink only
    Wire service; pages are not available to automated readers. Link-only.
  • axios.comlink only
    Blocks automated readers. Link-only.
  • bloomberg.comlink only
    Paywalled; blocks automated readers. Link-only.
  • economist.comlink only
    Paywalled; blocks automated readers. Link-only.
  • ft.comlink only
    Paywalled; blocks automated readers. Link-only.
  • nytimes.comlink only
    Paywalled; blocks automated readers. Link-only.
  • reuters.comlink only
    robots.txt disallows automated access to articles. Link-only.
  • washingtonpost.comlink only
    Paywalled; blocks automated readers. Link-only.
  • wsj.comlink only
    Paywalled; blocks automated readers. Link-only.

How stories are chosen

In short: each pass reads a rotating batch of sources, drops anything unreadable, duplicated or too short, matches what is left to a fixed list of topics, and publishes one story. Which sources are read, the topic list and every score used are published, with the live settings, on How we choose, including where tone (VADER sentiment) is, and is not, part of the pick on this site. The feed itself is ordered newest first.

Images

A story's picture is the image its publisher chose for it. We keep our own re-sized copy so the page loads fast and does not depend on the publisher's server (copies are deleted after 9 days). Each such image carries a caption naming the site it came from, and the story links to the publisher's page. If you are the rights holder of a picture and want it taken down, or want your site shown with our neutral image from now on, write to us (below); we act on a request by removing the copy and recording your preference in the terms file above.

For publishers: contact and opt-out

You can tell us what you want in any of these ways, and we will follow it: a robots.txt rule for our bot name (HuntaegisBot); a noai / nosnippet robots meta tag; or an email to ross@arc-codex.com naming your site and what you would like (stop reading it, link only, no images, a correction). A request by email is recorded in the terms file above and applied without waiting for a reply from you.

We would rather hear from you than guess. If a credit, link or excerpt on this site is wrong, tell us the story's address and we will fix it.

Licences publishers declare

Some feeds state a licence in their own metadata. Where a feed declares a Creative Commons or similar licence we list it here; we do not infer a licence for anyone who has not stated one. Of the 311 hosts surveyed, 13 declare some copyright or licence statement in their feed.

  • www.cert.br: CC BY-NC-ND 4.0

This page as data