Image: substackcdn.com · rights & removal
How Claude Watermarks AI
Reporting by Sebastian Raschka AI MagazineRead the original at magazine.sebastianraschka.com
Executive Summary
Facts Only
* Anthropic announced watermarking for Claude model text outputs.
* The goal of watermarking is to allow authorized parties to identify AI-generated text.
* Text generation involves converting input to token IDs and obtaining a score distribution (logits) from the LLM.
* Token sampling selects the next token based on these scores, often involving softmax conversion and random sampling.
* Watermarking uses a secret key and previous tokens to derive a random seed for deterministic sampling.
* The watermarking is applied at the sampling stage, not inside the core LLM training.
* Detection requires access to the watermarking key and scoring functions.
* Watermark removal is challenging because watermark positions are not explicitly known.
* The detection method uses watermarking functions ($G1$ to $Gn$) and averaging scores over text segments.
Full Take
From the original · Sebastian Raschka AI Magazine
I recently posted a Substack note about Claude’s new watermarking process and implementation. Since it’s such a popular topic and sparked such a lively discussion, I thought it might be interesting to go into a bit more detail when explaining how it works.Read the full story at magazine.sebastianraschka.com
Sentinel — Human
This text exhibits strong human authorship characterized by a highly engaging, pedagogical narrative style that weaves technical detail with personal reflection, though it relies on synthesizing external research.
