Home About Research
97.7MParameters - Alma 2 Base
16kLocal Alma 2 tokenizer
600kCompleted optimizer steps
19.66BCompleted token presentations
BaseCurrent Alma 2 stage
90/90Founder answers approved
6/6Audit rulings resolved
0.1.0Approved governing specification
0Constitutional fine-tune runs
Not yetPromotion-grade evaluation
Alma 1 identity and constitutional routing proof
Alma 2 Base 97.7M from-scratch general-language pretraining complete
Alma 2 Constitutional governing specification approved; training has not started

A language foundation first. A constitutionally trained Noema only after proof.

Alma 1 proved the constitutional routing loop at 59 million parameters: 12 layers, 8 attention heads, width 640, and a 128-token context. Alma 2 is a separate 97.7-million-parameter base with 12 layers, 12 heads, width 768, a 512-token context, and a locally trained 16k byte-level BPE tokenizer.

Alma 2 currently knows statistical language patterns, not a reliable account of its own identity, creator, authority hierarchy, tool limits, or response style. Those behaviors are being specified now and must be taught through a separately reviewed curriculum, then tested on frozen evaluations before promotion.

Behavior is specified, reconciled, taught, and then tested.

Alma 1's constitutional-routing proof, measured.

6.0 3.0 0.0 0 1000 2000 training iteration train loss validation loss

Loss falls from 5.64 (a model guessing blindly over 256 tokens) to 0.08. The deck is memorized early; the tail sharpens the hardest cases.

37k 18k 0 successive training runs →

The corpus is grown deliberately — ~20k to ~37k characters across runs — adding one focused brain part at a time, never a noisy pile.

One Alma 1 2,000-iteration constitutional-corpus run finished in 179 seconds on a single GPU. The long Alma 2 base campaign is a separate workload. Training inputs are manifest- or corpus-bound, and material runs retain their configuration and result.

Two Alma 1 evals, with failures kept in the record.

Foundation eval

A fixed scoreboard of prompts checks that each category routes correctly — greetings to greetings, identity to identity, unknowns to honest limits. Deterministic pass/fail, no model judging another model.

Full-corpus eval

Alma 1's production checkpoint is tested against the exact 448 examples it learned; it reproduced 433. Its historical reviewed corpus later grew to 478 examples. That older deck is not the new Constitution curriculum; conflict detection remains a gate for any future examples.

Verified Alma 1 foundation score: 69 / 69 (100%)

Every category routes correctly. A shorter 2,000-iteration run had left four short-prompt collisions — identity ("What are you?"), ownership privacy ("Who owns Mardenic?"), and two boundary wordings ("Are you human? / alive?"). Running the full training cycle resolved all four.

Material pass, fail, and exploratory results remain in the Training Log, with the model stage identified so Alma 1 and Alma 2 measurements are not confused.

Alma 2 base pretraining is complete. Constitutional work is next.

The 97.7M-parameter base completed 600,000 steps and 19.66B token presentations, and a protected weights-only copy passed strict reload checks. The founder interview is complete at 90/90, all six contradiction rulings are resolved, and the four governing artifacts are founder-approved as specification 0.1.0. Curriculum creation, constitutional fine-tuning, and promotion-grade behavioral evaluation have not started.

Read the training path →  ·  Review the completed Alma 2 run →

A controlled search tool, not unrestricted internet access.

What works now

Noema-SearchBot combines a focused Mardenic-controlled index with scoped web search and fetch. A limited host-mediated bridge smoke test completed one search, one generation, and a traceable citation.

What remains disabled

Alma 2 cannot autonomously browse or call tools. The host chooses the retrieval boundary, retrieved instructions remain untrusted, and future model-directed use must pass constitutional and security gates.

Open about the work. Closed about the keys.

What we publish

Noema and the Alma model line, the architecture, the operating loop, the eval methods, and the numbers — including the failures.

What stays private

Credentials and infrastructure, the raw training files and internal prompts, private operators, and unreleased implementation details.