Skip to content
Infrastructure of Meaning

Developer Documentation

The architecture and design choices behind Arc Codex.

Flask + Next.jsRedis + SolrLocal + cloud inferenceLocal speech synthesis
I

Stack Overview

Backend
Python 3.12.15 / Flask 3.1.3 / gunicorn 25.1.0
Frontend
Next.js 16.3.8 / React 19.2.4 / TypeScript 5
Database
Redis 8.10.0 (in-memory store + work streams)
Library DB
SQLite (public-domain book corpus)
Search
Apache Solr 9.9.0 (full-text)
AI Inference
Ollama — gemma4:e2b locally, gemma4:31b-cloud on escalation
Speech
Kokoro neural TTS — local synthesis, MP3 output
Metrics
Prometheus + Grafana 10.2.7 (corpus, pipeline and host telemetry)
Auth
Auth.js v5.0.0-beta.32 — Google and GitHub OAuth, JWT sessions
Proxy
Caddy (automatic TLS via Let's Encrypt)
Process Mgr
a control script + systemd (auto-starts on boot), with a watchdog

Read from the running system on 2026-10-08

II

Where Things Run

Arc Codex is not a single box. Work is split across 3 machines with distinct jobs, so that a heavy job on one never makes another wait. These are roles, not addresses: what each machine is for, read from the fleet registry.

resolute — Web and data host
Serves the sites and holds the data: reverse proxy, Redis, Solr, the web backends and frontends, the ingest and analysis workers, and the small council models.
warden — Hunt inference, narration, media backend, backup disk
Runs Huntaegis's analysis model, Arc's translation and broadcast-script model, the shared narrator (the only machine that synthesizes speech), the video backend, and the backup disk.
spectre — Arc inference
Runs Arc's analysis model (primary and fallback).

Speech is synthesised in exactly one place, by the shared narrator, one job at a time. It shares a machine with an analysis model, so it runs at the lowest CPU priority and gives way to analysis whenever both want the machine. No other service produces speech.

III

Services & Supervision

Within each machine, the services are managed by a single control script and auto-start on boot. A watchdog supervises them at runtime and restarts any that crash, distinguishing a deliberately-stopped service from a failed one.

What runs: RSS ingestion (the Scribe); the analysis worker (the Analyzer) that produces the three A.R.C. passes; publishing of user submissions; automated posting to Bluesky, Mastodon and Facebook (each toggleable at runtime with no restart); an email digest; the shared narrator; and the Next.js frontend. Telegram support exists in the code but is switched off; nothing is posted there.

Posting is fail-safe, not fire-and-forget: the posters track which articles they have published separately, pace themselves, and wait for the counter-analyst comment, so a mid-publish failure re-tries on the next cycle rather than double-posting or silently dropping.

IV

Reverse Proxy

A Caddy reverse proxy terminates TLS and routes requests: authentication and user-preference calls are handled by the Next.js application server, and the remaining API traffic goes to the Flask backend. Everything else renders from Next.js. Automatic certificate management is handled by Caddy.

V

Public API

Arc Codex exposes a read-only public HTTP API over the same domain. It serves the article feed, individual articles with their full analysis, full-text search, an RSS feed, the wiki directive pages, the public-domain library, and the machine-readable sitemap. Signed-in users can additionally submit content for processing and manage their own preferences.

Feed stories carry a short excerpt of the source text rather than the full text: the analysis is ours, the article belongs to its publisher, and the page links to the original. Machine-readable discovery surfaces — /rss.xml, /sitemap.xml, /news-sitemap.xml, and /opensearch.xml — are published and kept current automatically. A removed or never-existing article answers 410 Gone with a no-index header, so search engines stop asking.

VI

Data Model

Live application data is held in Redis for speed; the public-domain book corpus behind the Library lives in SQLite; and full-text search is served from Solr. At a conceptual level the system stores:

Articlesexpand
Each article carries its source text, metadata, editorial directive, a reading-difficulty score, the three A.R.C. analyses, and an AI-content verdict. Articles are typed as either rolling news or durable reference content. Published news is retained for 30 days and then pruned, together with its narration.
Comments & reactionsexpand
Reader comments and per-comment reaction counts. The adversarial Counter-Analyst comment is a first-class, distinctly-styled entry.
Translationsexpand
Per-article, per-language translations are cached (a day, or a week for translations into English) so a repeat request is instant. Non-English stories are translated to English once, at publication, and keep a record of their original language.
Work queuesexpand
Analysis is handed between processes on a length-capped stream rather than an unbounded list. A burst of ingest cannot grow the backlog without limit, and a consumer that restarts resumes from where it stopped instead of replaying the corpus. The cap is deliberate: dropping the oldest pending work is preferable to exhausting memory on a machine that is also serving readers.
Accountsexpand
A minimal profile per signed-in user — identity from the OAuth provider, a preferred language and a default visibility for their own submissions. Authentication is stateless (JWT); no server-side session store is required.
VII

Authentication

Soft auth — the site is fully public. Signing in with Google or GitHub is optional and unlocks preferences, publishing, and private articles. There is no username/password fallback and no third-party tracking.

Sessions are JWT-based (stateless), and preference writes are accepted only from the application server itself — never directly from the public internet — so a user can only ever change their own settings. A private article answers as not found to everyone but its owner.

VIII

How a Story Is Chosen

Every 37 minutes the Scribe reads the next 100 sources from a rotating list, drops candidates that fail simple filters (duplicates, paywalls, too-short pages, unreadable text), scores the rest, and publishes the single best match against a ranked list of topics. The complete list of every score, filter and threshold, with the settings in force today and what each one cannot see, is on its own page.

These measure tone and language, not truth. There is no automated fact-check and no source-reliability score, and the page says exactly which scores take part in the choice and which only annotate it.

IX

Machine-Written Provenance

Some stories are written by our own generation, from a prompt a signed-in reader submits. Each is recorded as machine-written (authorship MWX) together with the model that wrote it and when; the account that prompted it is stored privately and is never part of a public response.

Because the authorship is known rather than estimated, such a story gets no Sentinel probability: Sentinel is a guess about text of unknown origin, and here there is nothing to guess. An upload that merely claims a human wrote it still goes through Sentinel; only our own generation is treated as definitive.

X

AI Pipeline

Inference is tiered and demand-gated: a compact local model (gemma4:e2b), running on the site’s own inference machine, handles the bulk of the work, and a larger cloud model (gemma4:31b-cloud) is reached only on escalation, within a monthly budget that resets on day 7 of the month. One counter covers every site on the account; when the allowance is spent, or the provider says to wait, work falls back to the local model or parks.

Analysis is queued as soon as a story is published, behind a reader’s own request: if someone opens a story before its analysis is done, it jumps the queue. Inference cost therefore tracks both ingest and readership. Translation degrades gracefully when a model is unavailable: “model unavailable” is shown rather than a hard failure.

In the feed, translation is a click, not an auto-fire. A scrolled feed holds many mounted cards; firing translation on each mount would overwhelm the inference tier. On a single article page, a reader’s preferred language does translate automatically, one article at a time.

XI

Audio & Narration

The audio a listener hears is not the source article read aloud. After the Red / Blue / Purple passes complete, an additional pass writes an original short broadcast piece from what those three found — the verified facts, the balanced summary, and the anti-pattern reading. That written piece, not the article the Scribe ingested, is what gets spoken. The reasoning is editorial: reading source prose aloud reproduces its framing verbatim; narrating from the site’s own analysis yields a piece whose voice is the site’s, whose claims trace back to what the analysis concluded, and whose length is chosen for the ear rather than the page.

The piece is handed to a neural text-to-speech model (Kokoro) which renders it locally — no cloud speech service, no per-character billing, and no third party receives the text. Long pieces are split into chunks, synthesised in sequence, then concatenated and encoded to a compact mono MP3 sized for slow connections. For security reporting the written piece is also normalised for speech: CVE identifiers and version numbers are spoken, while commands, file hashes and address lists are not read out.

One shared narrator for every site: one speech job at a time, low CPU priority, Huntaegis stories first (a daily cap), Arc stories fill the rest; mp3s are pushed to the serving host and verified before they are published, then deleted from the narrator. Publishing never waits on audio: the narrator picks the newest eligible story and a deferred one is simply retried.

Synthesis yields to analysis. It runs at the lowest CPU priority, a pre-flight check confirms there is genuine memory headroom before a run starts, and during the peak-hours pause (next section) it stops entirely. Audio is the part of the system that can afford to be late.

XII

The Peak-Hours Pause

The machines are shared with the people who run them. During the hours the operator is at work the heavy background work stops, so the machines are free for other use: ingestion, analysis, narration and catch-up jobs pause, and resume when the window ends. Reading the site, signing in and submitting a story are never paused (a submitted story is published at once; its analysis waits for the window to end).

May to September
weekdays 2 pm to 7 pm
October to April
weekdays 5 pm to 9 pm

Weekends are never paused. Times are America/Denver; the window changes with the season by itself, from the configuration rather than from anyone remembering to.

XIII

Observability

The pipeline is instrumented rather than trusted. Metrics are scraped continuously and rendered as dashboards covering ingest rate, analysis latency, inference tiering, and corpus-level qualities — the average reading difficulty and objectivity of what has actually been published, not merely how much of it there is. Host health (CPU, memory, disk) is collected from 3 of 3 machines right now.

Alerting distinguishes liveness from output. A worker publishes a heartbeat on a short expiry, so its silence is itself the signal; that is a separate question from whether the day produced many articles or few. Conflating the two produces an alarm that fires on every quiet afternoon and is therefore ignored when it matters.

Narration adds a third kind of check, because its failure mode is silent. A synthesis worker can be up and consuming memory while producing no audio at all — a heartbeat would still be green. The narration-liveness check therefore ignores process state and watches the output itself: if audio has not landed for the newest publishable articles within the expected window, that absence is the alert. Liveness is measured by what arrived, not by what is running.

An alert that cannot clear is not an alert. Conditions are edge-triggered and paired with an explicit all-clear, so a fault that resolves itself says so. Without that, a recovered incident and an ongoing one look identical from the outside.

XIV

Frontend Notes

  • Feed rendering. The lazy-loading feed structure is load-bearing — changes are surgical, never structural.
  • Theme layer. A single stylesheet layer is the source of truth for colours and overrides everything else.
  • Preferences. One context is the single source of truth for user preferences across the app.
  • App Router. Next.js 16 App Router with Turbopack. Not the pages router.
  • One card, two sites. The story card and several pages are the same file on both sites, enforced by a test; site differences live in configuration.
  • No ads. Fully ad-free by design. No ad networks, no analytics beacons.
XV

Search

Full-text search is served by Apache Solr 9.9.0, indexed over the article corpus (title, content, source, directive, and the reading-difficulty score). Search reconnects lazily so a restart of either the search engine or the application resolves itself without manual intervention.

XVI

Planned Features

Future roadmap
  • Topic / category preferences per user (today a user can set a preferred language and a default visibility; nothing more).
  • Near-duplicate detection (SimHash / MinHash). Exact duplicates are already skipped; a reworded copy of the same story is not.

Nothing in this list is live.

XVII

A Network, Not a Feed

Direction · Not yet built

Today, Arc Codex is one instance reading the world alone. The RSS ingestion is cheap; the analysis is the expensive part — the three A.R.C. passes, the sentinel verdict, the counter-analyst comment — and every instance that runs the pipeline re-derives conclusions its peers have already computed.

The direction is a network of instances that share their analyses rather than each re-computing them. An operator standing up their own node would contribute what their instance analyzes and draw on what others have already analyzed — a corpus larger than any one machine’s ingest, with cross-verification implicit in the moments several independent nodes reach the same read on the same story.

Trust would be per-node. An operator decides which peers they read; a peer’s consensus becomes visible without becoming binding — annotation, never authority, consistent with the editorial principle applied to the feed itself.

Nothing here is live. The pieces that would need to land first — a stable identifier for each analysis independent of the article’s local ID, a signature scheme so a shared analysis carries its provenance, and a peer-discovery layer that survives an operator leaving — are named to give the direction a shape, not to promise a date.

XVIII

A Cluster from a Cold Boot

Direction · Not yet built

An operator with a spare PC on the same LAN as an instance should be able to boot it from a USB stick and have a new analyst node come up reachable, cloud-off, and ready to be named in arc.cfg. No manual OS install, no per-machine configuration, no key exchange after the fact — the appliance is the whole setup.

Two of the pieces exist. A bootable image brings a cold machine up to a reachable inference host in one boot — it carries Ollama, the local model, network config, and a systemd unit that binds the service to the LAN behind the host firewall. An Ansible role covers the same territory for machines that already have an operating system installed. Both are working inputs to a cluster; neither is a cluster on their own.

What isn’t there yet is the connective tissue that would turn one working node into a member of a cluster: a way for a booted node to announce itself to an instance rather than the operator hand-copying an address; a nodes section in arc.cfg that enumerates the fleet; and dispatch inside the Analyzer that spreads work across the named nodes instead of pinning to a single OLLAMA_URL. The image and the role that already exist live in private trees while a secrets split completes — the built USB currently bakes an operator SSH key, and the Ansible tree names hosts by their LAN address — so publication follows that work, not this section.

Nothing here is live. The pieces that would need to land first — node self-announcement, a nodes section in arc.cfg, a multi-analyst dispatcher inside the Analyzer, and the secrets split that makes the image and the role safe to ship — are named to give the direction a shape, not to promise a date.

© 2026 Arc Codex

github.com/hapnesbitt/arc-codex

Harold Edwin Ross Nesbitt III

Fort Collins, CO · 40.5853° N, 105.0844° W

A.R.C. Framework v7.38 · Connection Secure