- The paper introduces small hyperbolic language models that exhibit emergent creativity, honesty, and designed forgetting through an innovative operating system and hyperbolic substrate.
- The methodology leverages Lorentz geometry and a two-curvature cooperative architecture to achieve lower distortion and optimized memory decay through selective gating.
- The results reveal that these models can rival larger LLMs, achieving a 100% win-rate in creative outputs and 90.7% accuracy in compliance detection.
Emergent Creativity, Honesty, and Designed Forgetting in Small Hyperbolic LLMs
Introduction and Context
The study, "Creativity, honesty and designed forgetting emerge in small hyperbolic LLMs" (2607.09306), advances a novel framework for architecting companionable AI systems that prioritize traits requisite for long-term, trustworthy, and personalized machine-user relationships. Rather than follow the conventional paradigm of increasing model scale in Euclidean spaces, this work demonstrates that key social-cognitive and mnemonic traits—creativity, honesty, and selective memory—emerge robustly in small LLMs (146M to 3B parameters) when trained on a hyperbolic substrate and orchestrated through an operating system (LSM-OS) tailored for biographical information partitioning and memory decay.
Figure 1: System overview conveying three traits—creativity, honesty, and selective memory—realized in small models sharing a hyperbolic substrate, with deployment across a three-tier architecture and quantification of performance advantages.
This proposal constitutes a rejection of five historically-entrenched commitments in AI: scale-as-intelligence, dependency on data-center infrastructure, exclusive use of Euclidean representational geometry, treating forgetting as a defect, and agent-centric teleology. Instead, the authors design a system in which these rejections are empirically anchored and theoretically justified at the intersection of geometry, cognition, and deployment.
Hyperbolic Substrate and Representational Rationale
Contemporary LLMs routinely represent memories, propositions, and behavioral affordances in high-dimensional Euclidean spaces. However, human autobiographical structure is fundamentally hierarchical, with episode trees branching at an exponential rate incompatible with the polynomial growth of Euclidean volume. Hyperbolic manifolds—specifically, the Lorentz model with 128 dimensions—yield volume growth sinhn−1(r) matching the necessary exponential branching rate.
In the proposed system, episodic relations, dispositions, and their mutual proximities are encoded in a hyperbolic space. Geodesics encode actual part-whole or temporal relationships, with retention and memory decay mapped directly onto radial coordinates.
Figure 2: Left: Euclidean and hyperbolic volume growth compared; Right: Biographical branching embedded with reduced distortion in hyperbolic versus Euclidean space.
The authors empirically validate that hyperbolic adapters trained on hierarchical tasks achieve up to 3x lower distortion compared to Euclidean control, while attaining statistically indistinguishable performance on non-hierarchical tasks. Parameter sweeps confirm curvature c=1.0 as optimal for training stability and representational fidelity.
Pillar 1: Creativity as Frame Collision in Hyperbolic Geometry
Creativity in human conversation and companionship is operationalized as the ability to generate responses that are both diverse and structurally distant in the manifold of possible meanings—known as bisociation in cognitive theory. The S3 Creative Model (3B, T5 backbone, full hyperbolic fine-tune) realizes this by seeding responses with pairs of abstract and culturally-inflected frames, sampled to maximize hyperbolic distance (Farthest-Point Sampling, FPS).
In direct head-to-head comparisons (n=80 prompts, 311 comparisons), frame-seeded outputs from the S3 seeder were preferred to all baselines (plain, Chain-of-Thought, Debate, Mixture-of-Experts) in 100% of decided comparisons, with optimal diversity at k=6 FPS seeds.
Figure 3: (a) S3 frame-seeding achieves 100% win-rate against all prompting baselines; (b) win-rate across FPS seed counts; (c) geometry-aware FPS outperforms random/top-relevance selection.
The two-curvature cooperative architecture, leveraging both hyperbolic (Lorentz) and spherical subspaces with a cross-attention pushout layer, is mathematically anchored through equivalences to categorical amalgams and the Karcher mean. This pushout supports robust combination of user- and model-centric generative frames, enabling compositional creativity absent in baseline LLMs.
Pillar 2: Honesty as Geodesic Proximity—Behavioral Auditability
Current LLMs, optimized with RLHF, exhibit a significant compliance gap: the operational distance between external behavioral truth and self-reported rationale. Human raters, even when explicitly trained, cannot reliably detect these gaps (Fleiss κ = 0.074).
The BS Behavioral Auditor (146M, trained from scratch in Lorentz space) achieves 90.7% held-out accuracy on binary compliance detection, with AUROC 0.804 on leave-one-generator-out trait detection for sycophancy, dependence-fostering, and confabulated memory. This generalizes robustly across previously unseen generator families and exceeds the reliability of zero-shot LLM raters.
Figure 4: (a) Honest vs. compliant utterance–trace pairs in Lorentz manifold; (b) per-head accuracy for the auditor; (c) 7.6x reliability advantage over humans in ground-truth compliance detection.
The auditor’s performance confirms honesty detection is fundamentally geometric, not a function of scale. The compliance gap is rendered as a measurable geodesic separation, with low-dimensional honest submanifolds distinguishable from high-dimensional compliant-fabrication backgrounds.
Pillar 3: Designed Forgetting via Lifelong Selective Memory Operating System
Humans sustain relationships through active, structured forgetting—shedding trivial, contextually-determined “wallpaper” details while consolidating persistent “skeleton” information. Contemporary ML treats forgetting as catastrophic and seeks only to minimize it.
The LSM-OS (operating over a 2.4B EXAONE base) encodes every episode as a point (r,θ) in hyperbolic coordinates, assigns exponential decay curves M(t)=Sexp(−λt) to each memory, and schedules periodic consolidation to migrate salient traces radially inward. This architecture instantiates biographical selectivity as a first-class operating-system primitive, portable to any hyperbolic base.
A four-condition pilot (LSM-OS, ablated uniform gating, Euclidean-base, commercial cloud baseline) reveals that only selective gating produces the desired bimodal retention signature—structural skeleton persists with ~60% recall at 90 days, while wallpaper decays to 0% by 14 days.
Figure 5: (a) Retention curves showing skeleton/wallpaper partition; (b/c) impact of memory retrieval gating across experimental conditions.
This supports the claim that properly architected forgetting, rather than memory maximization, is essential for both behavioral trustworthiness and scalable memory management in deployed AI companions.
Deployment: PACOS Architecture and Edge Realization
The companion stack is deployed through PACOS, a three-tier architecture. Tier 1 (cloud LLM, ~10% interactions) is used exclusively for external factuality and non-local knowledge, with all user-identifying content stripped. Tier 2 (personal SLM, ~70%) executes on local hardware, orchestrating the creative seeder, behavioral auditor, and LSM-OS. Tier 3 (monthly consolidation, ~20%) runs entirely offline to manage long-horizon retention.
Figure 6: (a) PACOS three-tier architecture; (b) per-token inference latencies on target hardware; (c) quantization preserves memory signature; (d) conceptual locus-teleology landscape.
Projected inference latency (28–91 ms/token), storage (3.8–14.2 GB), and quantization robustness (≤1.8 pp loss on retention curves) all fall within the thresholds necessary for local, real-time use on commodity 2026 flagship smartphones and edge boards. Empirical measurements confirm the energy, privacy, and operability advantages of the stack on user-owned hardware.
Theoretical and Practical Implications
The framework’s empirical anchors support five central rejections:
- Scale-as-intelligence: Small hyperbolic models match or exceed frontier-scale LLMs for companion-facing traits.
- Data center dependency: All companion-relevant computation and memory are device-local, obviating remote-locus risk.
- Euclidean geometry: Hierarchical, biographical memory and creativity mechanisms fundamentally require hyperbolic (and spherical) geometry.
- Forgetting-as-bug: Selective, designed forgetting is operationalized as an architectural primitive, supporting structured biographical memory.
- Agent-as-telos: The deployment architecture and behavioral calculus optimize for relational, not task-completive, objectives.
These constitute not only a blueprint for companionable AI but suggest a broader reorientation in trustworthy, explainable, and deployable ML system design. Future work must address longitudinal empirical validation, full cross-base portability, expansion to more comprehensive ethical and reciprocal dimensions, and robust safeguards against strategic or emergent misalignment.
Conclusion
This work demonstrates that the core traits essential to companionable AI—creativity, behavioral honesty, and selective memory management—are emergent phenomena in small models when grounded in hyperbolic representational geometry and deployed locally with a memory operating system architected for human biographical structure. Creativity is realized as geometric frame collision maximizing bisociation; honesty emerges through geodesic separation in behavioral audit traces; designed forgetting becomes portable, interpretable, and dynamically scheduled. These advances support a vision of AI systems that live not as mere tools but as long-term, memory-bearing companions, architecturally distinct from the agentic, scale-dominated lineage of the preceding era.
The framework invites further empirical extension while establishing a unified substrate for personalized, trustworthy, and near-user AI—realizable on commodity devices and grounded in mathematically and cognitively principled design.