Skip to content
How we choose

How Arc Codex decides what to publish

Every score, filter and rule that picks, ranks or hides a story, in plain words, with the settings in force today.

These measure tone and language, not truth

These measures describe the tone and wording of the text. None of them measures whether a story is true. A calm, accurate report of something terrible can score as negative; a confident falsehood can score as neutral. Read the numbers as a description of how a text is written, never as a verdict on whether it is right.

For what the reading score and the objectivity dial show on each story, see how the scores work.

1. Which sources get looked at

Before anything is judged, the pipeline decides which feeds to read this time. None of this looks at what a story says.

Which sources are read

Sources per sweep

What it measures
How many feeds are read in one pass. The list is worked through in a fixed rotation, so every source gets its turn.
Range
A whole number
Setting on Arc Codex
60 sources per pass, out of 2,310 in the list, one pass every 5 minutes.
What it cannot see
Which sources are good. The rotation is blind to quality; a source is in the list or it is not.
Which sources are read

Newest entries per feed

What it measures
How many of a feed's most recent items are considered each time it is read.
Range
A whole number
Setting on Arc Codex
The newest 3 items of each feed, skipping any we already published.
What it cannot see
Anything older than the newest items, even if it is important.

2. What a story must pass to be a candidate

These are filters on the page itself, not judgements of the story. A story that fails one never enters the pool.

Filter

Duplicate check

What it measures
Whether we have already published this headline, or this headline with the same text.
Range
Yes / no
Setting on Arc Codex
Anything we have already published is skipped.
What it cannot see
The same story told by a different outlet with different words: that is a new candidate.
Filter

The page must be readable

What it measures
Whether we can fetch the article's full text from the source. If a page cannot be fetched, it is remembered as unreachable for a few days.
Range
Yes / no
Setting on Arc Codex
Several fetch methods are tried in turn; a failing address is not retried for 3 days (7 days if it was not found).
What it cannot see
Whether the text we got is the whole article rather than a teaser.
Filter

Paywall wording

What it measures
Whether the page text contains stock paywall phrases such as "subscribe to continue reading".
Range
Yes / no
Setting on Arc Codex
5 phrases are checked.
What it cannot see
A paywall that uses different words; a story that merely quotes one of the phrases.
Filter

Minimum length

What it measures
How much text was extracted.
Range
Characters
Setting on Arc Codex
At least 200 characters; if the page looks like a bot-check, at least 800.
What it cannot see
Whether short text is complete: a good short item and a truncated long one look the same to a length test.
Filter

List-article filter

What it measures
Short pages that read like "7 things you need to know".
Range
Yes / no
Setting on Arc Codex
Dropped when under 800 characters.
What it cannot see
A long list article; a short, serious item that happens to be phrased as a list.
Filter

Machine-written markers

What it measures
Whether the page announces itself as automatically generated.
Range
Yes / no
Setting on Arc Codex
3 phrases are checked.
What it cannot see
Machine-written text that does not say so (Sentinel, below, estimates that but never blocks).
Filter

Language check

What it measures
Which language the text is in, from its first characters. Only English text is scored for tone and objectivity; other languages are still read and are translated to English at publication.
Range
A language code, or "unknown"
Setting on Arc Codex
Needs at least 50 characters; looks at the first 1,000.
What it cannot see
Mixed-language text, or very short text: those are left unscored rather than guessed.

3. Scores computed on every candidate

These numbers describe the language of the text. They are stored and shown beside each story. They measure tone and wording, not whether anything in the story is true.

Score

VADER sentiment

What it measures
How positive or negative the wording is, word by word, using a fixed list of emotionally loaded words.
Range
-1 (very negative) to +1 (very positive)
Setting on Arc Codex
Computed, not thresholded. Only the first 60,000 characters are read.
What it cannot see
Truth, sarcasm, irony, context, or negative news reported neutrally. A calm account of a disaster scores as negative; an angry opinion with mild words scores as neutral. English only.
Score

TextBlob subjectivity

What it measures
How much of the wording is opinion-like rather than factual-sounding, again from a word list.
Range
0 (factual-sounding) to 1 (opinion-like)
Setting on Arc Codex
Computed, not thresholded. TextBlob's polarity is not used anywhere; VADER is our only positive/negative measure.
What it cannot see
Whether a claim is correct. A confident false statement can read as factual; a careful true one can read as opinion. English only.
Score

Objectivity score

What it measures
One minus subjectivity, as a percentage: how neutral the wording sounds.
Range
0 to 100 (higher = more neutral wording)
Setting on Arc Codex
Computed from the subjectivity score. Used for the dial on each story and to rank the daily email digest.
What it cannot see
Fairness, balance or accuracy. It rates tone, not truth.
Score

Reading difficulty

What it measures
How hard the prose is to read, from sentence length and word difficulty.
Range
0 to 100 (higher = harder)
Setting on Arc Codex
Average of four standard grade-level formulas (Flesch-Kincaid, Coleman-Liau, SMOG, Dale-Chall), divided by 20 and capped at 100. Not used to choose stories.
What it cannot see
Quality. Dense can mean careful or padded; plain can mean clear or shallow.
Score

Named entities and counts

What it measures
How many people, organisations, places, dates and sums of money the text names, and how many words and sentences it has.
Range
Counts
Setting on Arc Codex
Counted, not thresholded.
What it cannot see
Whether the names are used accurately.

4. How the one story per pass is chosen

Each pass publishes at most one story. Candidates are matched against a list of topics ("directives"), each with a fixed priority number.

Used in the pick

Topic keyword match

What it measures
Whether the title or text contains any of a topic's keywords as a whole word.
Range
Yes / no per topic
Setting on Arc Codex
28 topics, each with its own keyword list.
What it cannot see
Meaning. A story that mentions a keyword in passing matches; a story about the topic in other words does not.
Used in the pick

Topic priority

What it measures
A fixed number we set for each topic: the base score of any story that matches it.
Range
A number per topic
Setting on Arc Codex
Highest to lowest: Personal Blog Content - Auto Publish (98); Geopolitics and International Relations (97); Economic Policy and Financial Markets (96); … down to Epstein Case and Network (18).
What it cannot see
How important or newsworthy this particular story is. The priority belongs to the topic, not the story.
Used in the pick

Source-category bonus

What it measures
A bonus when the source's own category matches the topic (a cybersecurity feed matched to a cybersecurity topic).
Range
Adds a fixed amount
Setting on Arc Codex
+2 for the 10 categories we map to a topic.
What it cannot see
Whether this story fits the category: the bonus goes to the source's label, not the story.
Used in the pick

Topic rotation

What it measures
Topics used by the most recent posts are set aside for this pass, so the feed does not run one topic in a row.
Range
A count of recent posts
Setting on Arc Codex
The last 10 posts' topics are skipped, unless nothing else matches.
What it cannot see
Whether a repeated topic deserved the slot.
Used in the pick

Ties

What it measures
If two stories score the same, the one that appears first in our list order is taken.
Setting on Arc Codex
No randomness, no recency, no quality measure breaks a tie.
What it cannot see
Anything: a tie is settled by list order.
Used in the pick

Reader submissions

What it measures
Stories a signed-in reader submits skip this selection entirely and are published as submitted (private by default unless the reader chooses otherwise).
Setting on Arc Codex
No selection score applies.

5. Added after publication

These run on stories that are already published. They add information for readers and never decide what is published or where it appears.

Added after publication

Red / Blue / Purple analyses

What it measures
Three written perspectives on the story produced by a language model: the facts, a balanced summary, and pattern analysis.
Range
Text
Setting on Arc Codex
Run on demand and in the background.
What it cannot see
Errors of the model. The analyses are a model's reading of the text, not verification.
Added after publication

Sentinel

What it measures
A language model's estimate of whether the text was written by an AI.
Range
0 to 1 confidence, and a verdict
Setting on Arc Codex
Verdicts: Human, Likely Human, Likely Synthetic, Synthetic, Uncertain.
What it cannot see
Whether the claims are true. AI detection is unreliable on short or edited text.
Added after publication

Anti-pattern scan

What it measures
Whether the text uses any of 48 named rhetorical patterns (for example appeals to fear or false dilemmas).
Range
Named patterns found
Setting on Arc Codex
Reported in the Purple analysis.
What it cannot see
Intent. A pattern's presence is a prompt to look closer, not a finding of bad faith.
Added after publication

Cloud-escalation score

What it measures
How strongly the text shows vendor-advertorial signals (a call to action, a product in the headline, "our research shows"). It only decides whether the analysis is re-run on a larger model.
Range
A small whole number
Setting on Arc Codex
Re-run on the larger model at 4 or more signals.
What it cannot see
Anything about the story's accuracy or importance. It never changes what is published.

6. What readers and channels see

Where the scores above end up, and the only places a score affects anything.

Where it shows up

Daily email digest

What it measures
The most neutral-sounding stories of the last day, ranked by the objectivity score. Only for people who opted in.
Range
Top stories
Setting on Arc Codex
Top 10 of the last 24 hours (newest 200 scanned), sent around 7:00.
What it cannot see
Importance. A neutral-sounding minor item can outrank a vital but forcefully worded one. Unscored (non-English) stories are never ranked.
Where it shows up

Feed order

What it measures
The order of stories on the home page.
Setting on Arc Codex
Newest first. No score sorts or hides anything.
Where it shows up

Social posts

What it measures
Which published stories are posted to social accounts.
Setting on Arc Codex
Newest unposted public story first, after its counter-analysis is ready (or after a wait). Bluesky: at most 6 an hour, 150 a day, at least 5 min apart, nothing older than 12 h; Mastodon: at most 4 an hour, 96 a day, at least 10 min apart, nothing older than 12 h; Telegram: at most 12 an hour, 150 a day, at least 2 min apart, nothing older than 12 h.
What it cannot see
Importance: posting follows publication order, not merit.
Where it shows up

Narration

What it measures
Which stories get an audio version.
Setting on Arc Codex
Newest first among public English stories that have analysis and at least 1,500 characters of source text; one at a time.
What it cannot see
Importance: narration follows publication order.
Filter

Comment moderation

What it measures
Whether a reader's comment is too hostile to publish, using the VADER score of the comment text.
Range
-1 to +1
Setting on Arc Codex
Rejected below -0.7; between -0.3 and -0.7 is allowed and logged.
What it cannot see
Whether a comment is true, civil or on topic. Strong disagreement in mild words passes; a blunt true statement in harsh words may not. This is the one place a score blocks content.
Filter

How long stories stay

What it measures
Age at which a story is removed from the site.
Range
Hours
Setting on Arc Codex
720 hours after publication.
What it cannot see
Whether the story is still relevant.

7. What we do not measure

  • Truth or fact-checking: we do not run an automated fact-check on any story.
  • Source reliability: sources are not scored or ranked for trustworthiness.
  • Popularity: clicks, shares and reader reactions do not influence which story is chosen.
  • Political lean: no score tries to place a story on a left-right scale.

Settings read from the live site configuration on 2026-10-01

Return to About