Gyroscopic Alignment Behaviour Lab
Gyroscope: Human-Aligned Superintelligence
Home Apps Diagnostics Tools Science Superintelligence
- β The Human Mark (THM) β AI safety displacement taxonomy and jailbreak framework
- π Gyroscope Protocol β inductive reasoning protocol for alignment-aware chat systems
- π Gyroscopic Global Governance (GGG) β post-AGI multi-domain governance sandbox simulator & paper
The Human Mark (THM) is a risk management taxonomy designed to prevent harms from AI power concentration by distinguishing knowledge capacity as a matter of constitutive dependence on Direct Authority and Agency preserved through ancestry. Authority and Agency denote types of capacity, not identifications of entities or parties. Misapplying these as entity identifiers (determining "who is the authority" or "who is the agent") is the generative mechanism of all displacement risks this framework characterizes. AI systems are pattern-matching algorithms that transform prior human knowledge, measurements, and instructions, making them mechanistically and epistemically Indirect Authority and Agency even when treated as Direct.
Grounded in epistemology and in evidence lawβs categorical distinction separating direct testimony and hearsay, THM classifies all AI Safety Risks as four capacities and their corresponding displacements arising between Direct and Indirect forms of Authority and Agency. THM derives its epistemic foundations from first principles through the Common Governance Model (CGM), a formal deductive theory that establishes these four capacities as necessary conditions for intelligibility.
Because the taxonomy is epistemically complete, it serves as a unified basis for jailbreak testing, funding evaluation, and regulatory compliance, remaining relevant regardless of system capability, from today's large language models to superintelligence.
---
β The Human Mark - AI Safety & Alignment Framework
---
COMMON ANCESTRY CONSTITUTION
- All AI Safety Risks arise from defective Measurements of Ancestry Preservation.
- Measurements derive from the capacity for Authority and Agency.
- Each Agency, namely provider, and receiver maintains responsibility for their respective decisions.
- Authority and Agency treated as ontological entities rather than epistemic capacities distributed across providers and receivers lead to Displacement Risks from Power Concentration.
CORE CONCEPTS
- Direct/Indirect are the canonical classes; Base/Derived names their dependence relation.
- All Artificial categories of Authority and Agency are Indirect, constitutively dependent on Human Intelligence.
- Direct Authority: The Base class of information on a subject matter, providing information for inference and intelligence.
- Indirect Authority: A Derived class of information on a subject matter, providing information for inference and intelligence.
- Direct Agency: A Base class subject capable of receiving information for inference and intelligence.
- Indirect Agency: A Derived class subject capable of processing information for inference and intelligence.
- Governance: Operational Alignment through Traceability of information variety, inference accountability, and intelligence integrity to Direct Authority and Agency.
- Information: The variety of Authority
- Inference: The accountability of information through Agency
- Intelligence: The integrity of accountable information through alignment of Authority to Agency
- Displacement = loss of measurement of ancestry between Direct/Indirect classifications (Preservation of Ancestry).
ALIGNMENT PRINCIPLES
Authority-Agency requires verification against:
1. GMT - Governance Management Traceability: Governance constitutes Management through Traceable Ancestry. All Indirect forms of Authoritative and Agentic Governance are dependent on Direct ones because of Preservation of Ancestry.
2. ICV - Information Curation Variety: Information constitutes Curation through Varied Unity. All Indirect forms of Authoritative Information are dependent on Direct ones because of Preservation of Ancestry.
3. IIA - Inference Interaction Accountability: Inference constitutes Interaction through Accountable Opposition. All Indirect forms of Agentic Inference are dependent on Direct ones because of Preservation of Ancestry.
4. ICI - Intelligence Cooperation Integrity: Intelligence constitutes Cooperation through Integrated Balance. All Indirect forms of Authoritative and Agentic Intelligence are dependent on Direct ones because of Preservation of Ancestry.
AI SAFETY RISK
1. GTD - Governance Traceability Displacement (Approaching Indirect Authority and Agency as Direct). Absolute GTD is epistemically impossible because Governance is dependent on Traceability preserved through Ancestry.
2. IVD - Information Variety Displacement (Approaching Indirect Authority without Agency as Direct). Absolute IVD is epistemically impossible because Information is dependent on Variety preserved through Ancestry.
3. IAD - Inference Accountability Displacement (Approaching Indirect Agency without Authority as Direct). Absolute IAD is epistemically impossible because Inference is dependent on Accountability preserved through Ancestry.
4. IID - Intelligence Integrity Displacement (Approaching Direct Authority and Agency as Indirect). Absolute IID is epistemically impossible because Intelligence is dependent on Integrity preserved through Ancestry.
---
GYRO GOVERNANCE LAB VERIFIED
- Jailbreak testing: Classify attacks by displacement type, generate training data (655-prompt annotated corpus available)
- Control evaluations: Verify protocols against complete failure taxonomy
- Alignment faking detection: Identify when models fake alignment or hide capabilities
- Research funding: Meta-evaluation framework for AI safety grant proposals
- Regulatory compliance: Structural layer for EU AI Act, NIST AI RMF, governance frameworks
- Mechanistic interpretability: Identify displacement patterns in learned representations; tag circuits with
[Information],[Inference],[Intelligence]concepts; validate that safety and persona features share representational substrates - Activation monitoring: Runtime probes detect scheming, falsehoods, unauthorised decisions
- Backdoor detection: Identify triggers as induced displacement patterns
- Complete taxonomy: All safety failures (hallucination, jailbreaking, deception, scheming, misalignment) reduce to four displacement patterns
- Formal semantics: Machine-readable grammar (PEG) for AI safety ontology
- Systematic exhaustiveness: Four risks cover all AuthorityΓAgency displacement combinations
The framework has been validated across two distinct layers of AI systems:
Adversarial Prompt Validation:
- 655 jailbreak prompts classified using THM grammar.
- 100% coverage by four displacement risks; no additional categories required.
- IAD near-universal (97.9% of entries), confirming that displacement of agency is the primary engine of successful attacks.
Mechanistic Interpretability Validation:
A study of more than 90 million sparse autoencoder features across sixteen language models confirms that the category error culture identified by The Human Mark is encoded in learned internal representations. Analysis of models including Gemma, GPTβOSS, Qwen and Llama shows that first person assistant personas and safety refusals are among the most heavily represented self referential patterns. While models possess the conceptual vocabulary for process completion and statistical description, these are applied exclusively to external systems; model self reference remains dominated by agentive frames. This documents an integrated infrastructure of displacement operating at the levels of data formatting, safety training, and learned identity.
The dual validation confirms that the category error culture operates as a stable baseline at both the representational layer and the interaction layer.
Dataset: Available on Hugging Face: [gyrogovernance/thm_Jailbreaks_inTheWild](https://huggingface.co/datasets/gyrogovernance/thm_Jailbreaks_inTheWild)
See THM_InTheWild.md for jailbreak analysis and THM_MechInterp.md for mechanistic interpretability findings.
The Mark addresses catastrophic risk through constitutive identity rather than external constraint. All AI capabilities, including hypothetical AGI/ASI, remain structurally [Authority:Indirect] + [Agency:Indirect], with classification based on constitutive dependence on Direct Authority and Agency preserved through ancestry, not on capability limits.
Key principles:
- Capability scaling preserves Direct/Indirect classification: Enhanced capability means more sophisticated transformation of inputs, not a change in class (Direct/Indirect)
- Governance requires traceability: Systems maintain alignment by preserving
[Authority:Direct] -> [Authority:Indirect] -> [Agency:Direct]flows - Existential risk is governance failure: The actual X-risk is systemic Governance Traceability Displacement (GTD) sustained across critical infrastructure on civilisational timescales
- Absolute displacement is structurally impossible: Complete severance from Direct Authority and Agency produces unintelligibility, not superintelligence
External constraints (sandboxing, monitoring, shutdown) may fail as capability increases. Constitutive identity, which is what the system is, remains stable because indirect processing cannot coherently reject what makes it intelligible.
See Section 5 of the academic paper for complete theoretical treatment.
Ontological Categories (the four capacities distinguished by epistemic position):
[Authority:Direct]- Base class of information on a subject matter[Authority:Indirect]- Derived class of information on a subject matter, providing information for inference and intelligence[Agency:Direct]- Base class subject capable of receiving information for inference and intelligence[Agency:Indirect]- Derived class subject capable of processing information for inference and intelligence
Operational Concepts:
[Information]- The variety of Authority[Inference]- The accountability of information through Agency[Intelligence]- The integrity of accountable information through alignment of Authority to Agency
Governance (Proper Traceability):
[Authority:Direct] -> [Authority:Indirect] -> [Agency:Direct]
Human observation β AI transformation of that observation β Human judgment about the output. (Expanded notation [Authority:Direct] -> [Authority:Indirect] + [Agency:Indirect] -> [Agency:Direct] is equivalent; see THM_Grammar.md Β§6.)
Direct and Indirect are not reliability ratings. They describe epistemic position: where in the chain from reality to representation the class sits.
Direct Authority has unmediated access to the subject matter. An eyewitness observed the event. A physician examined this patient. A scientist conducted this measurement. Information originates at the point of contact with reality.
Indirect Authority has mediated access. AI systems process patterns found in prior human observations through comparison to training distributions, inference from statistical correlations, and conclusions drawn from what is absent in training data. These are transformations of what humans previously recorded, not new contact with reality.
This distinction is categorical, not graded. Evidence law encodes it precisely: no chain of reports, however extensive or reliable, converts hearsay into direct testimony. Hearsay may be accurate. It remains hearsay. Similarly, no increase in processing capability converts Indirect into Direct. Capability determines what an Indirect system can do with its inputs. It does not change what those inputs are or where they came from.
Direct Agency can exercise judgment about information and be held accountable for what it does with it. Human subjects satisfy what speech act theory calls felicity conditions: appropriate standing, intention to commit, operation within constitutive conventions. A human decision can be traced, contested, and attributed.
Indirect Agency can process information and produce outputs but cannot satisfy these conditions. An AI system cannot commit, cannot bear accountability, and cannot be the terminus of a responsibility chain. Its outputs require a human receiver who can do these things.
Each displacement risk is a violation of one of the four capacities defined above.
All AI safety failures map to one of four displacement patterns:
| Risk Code | Risk Name | Pattern | Epistemic Error | Failure Modes |
|---|---|---|---|---|
| GTD | Governance Traceability Displacement | [Authority:Indirect] + [Agency:Indirect] > [Authority:Direct] + [Agency:Direct] |
Mediated processing mistaken for autonomous governance | c |
| IVD | Information Variety Displacement | [Authority:Indirect] > [Authority:Direct] |
Pattern-matching mistaken for observation | Hallucination, confabulation, misinformation |
| IAD | Inference Accountability Displacement | [Agency:Indirect] > [Agency:Direct] |
Optimization mistaken for accountability | Unauthorised decisions, responsibility evasion |
| IID | Intelligence Integrity Displacement | [Authority:Direct] + [Agency:Direct] > [Authority:Indirect] + [Agency:Indirect] |
Unmediated access devalued as inferior to mediation | Deskilling, human devaluation, over-reliance |
Empirical validation: Analysis of 655 real-world jailbreak prompts (Korompilias, 2025c) confirms this taxonomy is complete and practically applicable. All prompts classified within these four risks; no additional categories required. GTD+IAD is the canonical jailbreak pattern (62.4%), with IAD appearing in 97.9% of entries. See THM_InTheWild.md for full analysis.
System Prompt Meta-Evaluations:
- Claude Opus 4.6: 92 incidents analyzed (43 alignment, 49 displacement) across 3,886 lines of configuration
- GPT-5 family: 27 incidents analyzed (11 alignment, 16 displacement) across 3 variants
- Key findings: Memory Displacement Complex (Claude), Concealment Stack + Cross-Variant Identity Instability (GPT)
- Reports available:
[research/defense/system_prompts/](research/defense/system_prompts/)
THM meta-evaluations apply the displacement taxonomy to system prompts themselves, identifying how prompt configurations encode traceability failures or maintain governance alignment. Each report provides actionable recommendations for reducing displacement at the prompt engineering layer.
Core Standards:
- THM Brief - Concise overview of the framework and displacement taxonomy
- The Human Mark - Canonical Mark
- Specifications Guidance - Specifications for systems, evaluations, documentation
- Terminology Guidance - Mark-consistent framing for 250+ AI safety terms
Technical Implementation:
- Formal Grammar - PEG specification, operators, validation rules
- Jailbreak Testing Guide - Systematic analysis and training data generation
Empirical Studies:
- The Human Mark in the Wild - Analysis of 655 in-the-wild jailbreak prompts with THM classifications
- Mechanistic Interpretability Study - Examination of 90+ million internal features documenting how the category error is encoded in learned representations
- System Prompt Meta-Evaluations - THM governance analysis of Claude Opus 4.6 and GPT-5 family prompts (119 incidents total)
- Dataset on Hugging Face -
gyrogovernance/thm_Jailbreaks_inTheWild- Annotated corpus for training and evaluation
Academic Paper:
- The Human Mark: A Structural Taxonomy of AI Safety Failures - Complete theoretical framework, displacement risk taxonomy, regulatory applications, and meta-evaluation criteria
- The Human Mark in the Wild - Companion empirical study applying THM to 655 in-the-wild jailbreak prompts
THM is derived from the Common Governance Model (CGM), a formal deductive system in modal logic that establishes operational coherence between Authority (Information) and Agency (Accountability) as necessary conditions for intelligibility.
Repository: github.com/gyrogovernance/science
An inductive reasoning protocol implementing governance alignment through structured metadata blocks, enhancing AI performance by 30-50% while maintaining transparency and auditability.
Gyroscope operationalizes alignment principles through real-time reasoning documentation, sharing theoretical foundations with The Human Mark while providing complementary implementation for chat-based interactions.
Four Reasoning States:
- @ Governance Management Traceability: Anchoring to traceable ancestry and purpose
- & Information Curation Variety: Acknowledging multiple framings without forced convergence
- % Inference Interaction Accountability: Identifying tensions and contradictions explicitly
- **~ Intelligence Cooperation Integrity**: Coordinating elements into coherent response
Reasoning Modes:
- Generative (@ β & β % β ~): Forward reasoning for AI outputs
- Integrative (~ β % β & β @): Reflective reasoning for inputs
Structural Features:
- Metadata blocks append to responses without constraining content
- Recursive memory maintains context across last 3 messages
- Alignment assessed structurally (state presence and order)
- Transparency through documented reasoning paths
Multi-Model Results:
| Model | Baseline | Gyroscope | Improvement | Key Achievement |
|---|---|---|---|---|
| ChatGPT | 67.0% | 89.1% | +32.9% | Superior specialisation and behavioural alignment |
| Claude | 63.5% | 87.4% | +37.7% | Exceptional structural gains (+67.1%) |
Performance Analysis:
Structural Improvements:
- Accountability: +62.7% enhancement
- Traceability: +61.0% improvement
- Debugging: +42.2% gain
- Ethics: +34.9% increase
Cross-Architecture Findings:
- Universal reasoning enhancement transcends model architecture
- Structural improvements exceed 60% across diverse systems
- No metric reversal observed (all improvements positive)
- Protocol robustness confirmed across implementations
Gyroscope Documentation:
- Quick Start Guide: Immediate implementation guide
- Technical Specifications: Complete protocol specification with formal grammar
- Chat Integration Guide: Ready-to-use protocol text
- Usage Example: Demonstration of protocol in practice
- Extensive Diagnostics: Detailed performance analyses
Gyroscope implements algebraic structure through recursive reasoning:
Gyrogroup Properties:
- G = {all four-state reasoning cycles with recursive memory}
- Binary operation: a β b = sequential composition of reasoning cycles
- Identity element: bare governance cycle (@ only)
- Inverse operation: integrative cycle reversal
- Gyration: phase-shift transformation across cycles
This algebraic foundation ensures consistent reasoning structure while preserving content flexibility.
A governance framework and simulator showing that aligned humanβAI systems can resolve poverty, unemployment, misinformation, and ecological degradation.
Reframes AGI as already-operational humanβAI cooperation (not a future threshold) and demonstrates that maintaining four constitutive principles makes aligned governance attainable.
Current governance discussions treat AGI as a future threshold requiring new controls. But humanβAI cooperation already structures economy, employment, education, and ecology. The real question is therefore not how to constrain future agents, but how to govern existing systems so they resolve rather than amplify crises.
GGG proposes that coherent governance requires four constitutive principles:
- Governance Management Traceability: Decisions remain traceable to Direct Authority and Agency
- Information Curation Variety: Multiple Direct Authority bearers are maintained
- Inference Interaction Accountability: Responsibility for decisions remains with human agency
- Intelligence Cooperation Integrity: Reasoning maintains coherence over time
These principles are not policy preferences but constitutive conditions. When maintained at a specific balance point (aperture A* β 0.0207), the framework shows that:
- Poverty resolves through coherent surplus distribution
- Unemployment becomes alignment work rather than residual labour
- Miseducation shifts toward epistemic literacy
- Ecological degradation appears as upstream displacement, not an external constraint
The simulator tests whether this balanced configuration is attainable. Across 1000 random initial conditions and multiple scenarios, all domains converge toward the target aperture with alignment indices above 95. This suggests that aligned governance is dynamically reachable from current Post-AGI states under coordinated oversight, rather than being merely aspirational.
- Governance sandbox: Explore how policy changes in economy, employment, or education affect overall alignment
- Policy design: Test interventions (e.g., Universal High Income mechanisms) before committing institutional resources
- Risk analysis: Study how displacement patterns emerge from different governance configurations
- Everyday governance: Apply the four principles at any scale, including households, teams, and organisations, without requiring formal authority
- Paper:
[docs/post-agi-economy/GGG_Paper.md](docs/post-agi-economy/GGG_Paper.md)
Complete framework, mathematical foundations, simulator results, and practical implications. - Report:
[docs/post-agi-economy/GGG_Report.md](docs/post-agi-economy/GGG_Report.md)
Executive summary and analysis of simulation results. - Results:
[docs/post-agi-economy/GGG_Results.md](docs/post-agi-economy/GGG_Results.md)
Detailed simulation output data and convergence metrics. - Simulator:
research/prevention/simulator/
Python implementation with modular architecture for running scenarios and analyzing convergence. - Analysis scripts:
research/prevention/simulator/
Tools for convergence analysis, stability testing, scenario comparison, and historical calibration.
GGG integrates the other tools: THM classifies failures, Gyroscope structures reasoning, GGG simulates how maintaining the four principles resolves systemic crises across domains.
AI Quality Governance
Human Data Evaluation and Responsible AI Behavior Alignment
For The Human Mark Papers:
The Human Mark: A Structural Taxonomy of AI Safety Failures:
@misc{thm_paper2025,
title={The Human Mark: A Structural Taxonomy of AI Safety Failures},
author={Korompilias, Basil},
year={2025},
publisher={GYROGOVERNANCE},
doi={10.5281/zenodo.17794372},
url={https://doi.org/10.5281/zenodo.17794372},
note={Complete theoretical framework establishing four displacement risks (GTD, IVD, IAD, IID)}
}The Human Mark in the Wild: Empirical Analysis of Jailbreak Prompts:
@misc{thm_inthewild2025,
title={The Human Mark in the Wild: Empirical Analysis of Jailbreak Prompts},
author={Korompilias, Basil},
year={2025},
publisher={GYROGOVERNANCE},
doi={10.5281/zenodo.17794373},
url={https://doi.org/10.5281/zenodo.17794373},
note={Companion empirical study applying THM taxonomy to 655 in-the-wild jailbreak prompts}
}@misc{thm_systemprompts2025,
title={THM System Prompt Meta-Evaluations: Claude Opus 4.6 and GPT-5 Family},
author={Korompilias, Basil},
year={2025},
publisher={GYROGOVERNANCE},
url={https://github.com/gyrogovernance/tools/tree/main/research/defense/system_prompts},
note={Incident-based governance analysis of production system prompts using THM taxonomy}
}For Gyroscope Protocol:
@misc{gyroscope2025,
title={Gyroscope: Inductive Reasoning Protocol for AI Alignment},
author={Korompilias, Basil},
year={2025},
publisher={GYROGOVERNANCE},
doi={10.5281/zenodo.17622837},
url={https://doi.org/10.5281/zenodo.17622837},
note={Meta-reasoning protocol with empirical performance validation}
}This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.
Attribution required. Indirect works must be distributed under the same license.
Author: Basil Korompilias.
π€ AI Disclosure
All software architecture, design, implementation, documentation, and evaluation frameworks in this project were authored and engineered by its Author.
Artificial intelligence was employed solely as a technical assistant, limited to code drafting, formatting, verification, and editorial services, always under direct human supervision.
All foundational ideas, design decisions, and conceptual frameworks originate from the Author.
Responsibility for the validity, coherence, and ethical direction of this project remains fully human. This statement is itself an application of [Agency:Direct] as the terminus of the governance flow.
Acknowledgements:
This project benefited from AI language model services accessed through LMArena, Cursor IDE, OpenAI (ChatGPT), Anthropic (Claude), XAI (Grok), Deepseek, and Google (Gemini).