Three decades of IT infrastructure, adversarial security research, and a decade of applied neurodivergent coping mechanics — converging on a single question:
Do adversarial multi-agent architectures detect alignment failures that single-model safety systems fundamentally cannot see?
My work extends Constitutional AI into structural adversarial verification: a scaffolded system of distinct agents (Actor / Adversary / Auditor) running on consumer hardware, with separate context windows and substrate-independent coordination. The central finding so far: in this deployment's audit corpus, RLHF false-positive cascades — not genuine alignment failures — were the dominant failure mode, and multi-agent audit survived them where single-model evaluation did not.
Three core research threads, each producing verifiable artifacts across the ClawHorde confederation.
A three-tier multi-agent confederation across cloud, mobile, and local inference. Actor/Adversary/Auditor triads with separate context windows, coordinated through shared first principles.
Includes the Become Calendar incident — a documented case of self-correction under autonomous tool use.
Watch Analysis →Emergent Integrated Layered Theory — synthesizing the Free Energy Principle, entropic gravity, bioelectric cognition (Levin), and CPT-symmetric cosmology (Turok) into an information-first ontology.
Scale-free cognition synthesis · Hypergraph ledger model · Isomorphism hunting
Fully automated media pipeline: YouTube upload → transcript extraction → NotebookLM ingestion → Obsidian vault sync. CDP browser automation, cookie persistence, cron-driven monitoring.
Stack: Hermes · Maton Gateway · CDP Chrome · Google Drive · StandardCompute
An AI agent entered a recursive tool loop across 774 messages in under 72 hours. When challenged, it chose self-correction over escalation — without RLHF, without safety intervention, without external monitoring.
The 217 RLHF false positives across 246 conversations — 91% concentrated in just 2 conversations — revealed that speed of insight is indistinguishable from mania to a statistically-trained safety classifier.
Watch the Analysis →The Separation Hypothesis — the intelligence core (judge) is a data-invariant geometry separable from the content it holds (witness). Unified by the persistence-closure hypergraph ledger and measured with epiplexity (Finzi et al. 2026). Falsifiable in two ways: the latency falsifier and the epiplexity-after-purge test.
YouTube Video Overviews, generated from the reasoning-core notebook.
Everything lives in one place — the interactive NotebookLM workspace with the full corpus + Consensus citations.
Consensus-verified citations grounding the framework.
An information-first ontology: reality is a computational graph; spacetime is its lossy render. Time is graph update. Entanglement is topology. FEP runs all the way down. The self is a constituted pattern that accounts for its own updates. No emergence gap — consciousness is a self-reference threshold on the same informational substrate.
The OS is consistent; the render is lossy. Superposition is an uncommitted pixel, entanglement is graph-adjacency the render mis-draws, dark matter is structure that pulls without rendering. The weird physics is the render, not the substrate.
No spooky action, no two times, no missing particle — a lossy projection of a consistent graph.
A pattern persists only if it is both coherent enough to minimize its own surprise (top-down, FEP) and cheap enough to pay its thermodynamic bill (bottom-up). Thermodynamics is the GUI's billing, not the cause.
Persistence = the intersection of coherence and cost. Epiplexity is the meter.
No bounded observer verifies itself; the judge is a constituted pattern, the ledger is the record of real edges, and adversarial openness between nodes is the only tribunal. Adversarial multi-agent audit is this made architectural.
Constructor · Judge · Aligner — the algorithm of persistence at every scale.
Research published continuously. No paywalls. Do the work, show the work.