Home About Research

Learn language first. Earn Noema's behavior second.

Phase 1 — Complete
Build the language foundation

A fresh model, randomly initialized, completed 600,000 steps across 19.66B token presentations from quality-filtered public text. The protected base is preserved; its training-time monitoring loss is not clean held-out evidence.

Phase 2 — Specification approved
Teach and test Noema's behavior

The founder interview is complete at 90/90, the contradiction audit is fully resolved, and the four governing artifacts are founder-approved as specification 0.1.0. A separate reviewed curriculum will train a copy of Alma 2. Frozen evaluations must then test identity, truthfulness, safety, tool discipline, and capability retention.

No pretrained model weights or borrowed tokenizer are used. FineWeb-Edu remains third-party public source material. The founder workbook and future Constitution specify behavior; they do not become training data automatically.

Capability, earned one rung at a time.

01

Foundation — done

A small model that memorizes a reviewed deck perfectly: identity, honesty, boundaries, ownership. Foundation eval 100%. Proven, and the floor everything else stands on.

02

General language — base complete

Alma 2 completed its from-scratch base run and passed strict preservation and reload checks. Its small post-run diagnostic is not promotion evidence.

03

Constitutional behavior — specification approved

The governing specification is founder-approved. Next: author reviewed examples, freeze the behavior evaluation, fine-tune a copy, and run the proving loop.

04

Retrieval & tools — controlled prototype

Noema-SearchBot passed a limited host-mediated retrieval smoke test. Autonomous model-directed browsing remains disabled.

05

Alma 3 scale-up — proposed, not training

The 307.1M-parameter candidate waits on Alma 2 constitutional proof, data provenance and clean-split gates, mixture selection, and a separate shakedown.

Four days of training, preserved as evidence rather than a launch claim.

01

Checkpointed & recoverable

Atomic primary and previous checkpoints bounded interruption risk throughout the run. The completed source and a separate optimizer-free model copy are preserved.

02

Monitored throughout

The watchdog checked freshness, progress, loss, storage, and process health while training was active. Completion alerts and hourly monitoring are now off.

03

Final run record

The Alma 2 Run page preserves the completed progress bar, milestone history, final training loss, and contaminated-monitoring-loss warning.

04

Verified, not promoted

The completed weights passed strict reload checks. The small synthetic diagnostic is non-promotional, and constitutional behavior remains not evaluated.

Base pretraining and preservation are complete, and the Constitution is approved as governing specification 0.1.0. The current gates are the reviewed curriculum and the frozen behavior evaluation; only then does a new copy get fine-tuned. Read about the Constitution →

Honest about the trade, not just the win.

What we gain

A base that can be tested on unseen wording and broader tasks, rather than only the exact prompts used for Alma 1's constitutional-routing proof.

What it costs

A model that generates freely can also be confidently wrong. Clean holdouts, capability tests, safety probes, and an explicit promotion decision matter more than a falling training loss.