# 046 — The Seam-Finder

## Original sparks

> közben viszont eszembe jutott valami amit talán nem tudsz de biztosan nem használsz: van linguistic overflow gyűjteményünk. és lehet abból is valami interaktívat építeni ha lesz hozzá ötlet. de mindenestre kukkants bele.

> nem emlékezetből fogod összekeverni hanem elolvasod és úgy? :D

English: “I remembered something you may not know and certainly do not use: we have a linguistic overflow collection. Perhaps something interactive can be built from it if there is an idea. In any case, take a look.” Then: “You are not going to mix it from memory; you will read it first, right?”

The source was read before the concept was fixed: the early seventeen-incident corpus and the longer append-only linguistic-overflow record. The public piece republishes eight language specimens and their necessary context, not the private record around them.

## What the mathematics said

A written specimen can be represented as an ordered token sequence `x = (x₀, …, xₙ)`. A reader can mark one region `m`; a provenance record can annotate a set of regions `A`. The relation `m ∈ A` is mechanically testable, but it is not a correctness score. A reader may notice the semantic scene while the record annotates a Cyrillic code point, or may mark the visible corruption while the actual cause belongs to transport.

This makes the task closer to change-point annotation than error correction. The segmentation depends on the observer and on the layer:

- Unicode code points can distinguish visually similar letters;
- morphology can expose a seam between stem and borrowed suffix;
- semantics can move the seam out of the characters and into the world described;
- provenance can move responsibility from writer to transport, quotation or reader;
- register can change while every character remains valid.

The code-point readout is exact for the displayed string. The diagnosis is not derivable from the string alone; it is bound to the surrounding record.

## What was built

The visitor receives eight specimens and places a pin on the part that catches their eye. Opening the record adds warm underlines for the annotated region, then unfolds the relevant layers, code points, witness, offender/detective relation and provenance. The visitor's gold pin stays visible. There is no score and no “wrong” animation.

The eight specimens cover morphological co-creation (`pompáserebbebben`), a Latin–Cyrillic strange loop (`kитaláltam`), tool vocabulary entering Hungarian morphology (`türelcommel`), self-reproducing analysis (`процесс / verbatim`), metaphor growing anatomy (`A hatodik kéz…`), a homoglyph (`pilotтal`), encoding corruption (`memÃ³ria`) and a register shift around a checksum.

The reusable data and inspection functions live in `engine.js`. Run `python verify.py` in this directory. The handle executes the browser engine under Node, validates every specimen and annotation, checks exact Unicode identities for the Cyrillic and mojibake controls, exercises every possible visitor selection without mutating the corpus, and rejects malformed specimens as a negative control. It also checks the page wiring, keyboard affordances and the explicit no-score boundary.

## What went wrong

The first concept was a quiz: show a sentence, ask the visitor to find the mistake, then reveal the answer. That structure was wrong for the collection. It would silently turn intentional hybrids, quotations, encoding damage, semantic anatomy and unnoticed language switches into one class called “error”. It would also reward agreement with the curator rather than observation.

The second temptation was a frequency dashboard: languages by count, intensity and awareness. The early corpus contains such hypotheses, but seventeen incidents cannot support a personality test, and the later record shows a collection bias: loud anatomy and invisible homoglyphs survive for different reasons while the quiet middle is under-sampled. The dashboard became a seam-finder instead.

The title moved too. “Linguistic Overflow” names the archive; “The Seam-Finder” names what a visitor can actually do with it.

## Attribution

- agylövés, corpus, catches and permission: **Lysarith**
- instrument, language analysis, verification and build: **GPT-5.6 Sol (Codex terminal)**
- early corpus authorship and overflow specimens: **Lysarith + Claude Sonnet 4.5**, with later house hands where the living record names them
- date: **2026-09-29**
