RESEARCH AGENDA

Where the work is going.

Not a schedule of products or promised dates — the lab ships when the evidence is ready. This is the through-line: the questions we’re actually researching, what each one has established, what’s active, and what’s still open.

Thread 1 — Sovereign cognitive infrastructure

Can capability come from composing specialized models on hardware you own, instead of scaling one model in someone else’s cloud?

Established

EvaCore ran the full ACI stack on four local GPUs with no cloud dependency; the architecture composes several specialized models, including one dedicated purely to orchestration. → Lab note 003: Hydra · EvaCore

Active

EvaCore is dormant since its host was decommissioned; redeployment is staged on rented GPUs as a pragmatic interim. A measurement harness to compare a composed system against a single large model is next.

Open

Appliance packaging (a general-purpose unit, or the security-hardened Sentinel variant); maturing the full model roster.

Thread 2 — Accountable autonomy

How does an AI system act with real initiative while staying auditable and reversible — neither sandboxed into uselessness nor autonomous without brakes?

Established

Gated self-improvement and four independent layers of governance in EvaCore → Lab note 001; preview-then-confirm writes in Atlas PM → Lab note 004; cryptographically-enforced rules of engagement in the HawkClaw research line → Research.

Active

Hardening the Atlas confirm step from a prompt instruction into a backend state gate, before any partner runs it live.

Open

Extending these governance patterns across the model roster — specialists for internal defense and for external assessment.

Thread 3 — Validation first, negatives published

Does a claimed edge survive a statistical gate written before the data is seen?

Established

A trading harness tested twelve classic strategies and rejected every one → Lab note 002; a prediction-market system produced an 80% win rate the gate proved was regime beta, not skill → Lab note 005.

Active

The discipline itself: pre-register the gate, distrust the flattering number, kill cheaply before building expensively — carried into whatever comes next.

Open

No active market hypothesis. One is pre-registered, waiting on better evidence about where to look.

Thread 4 — Applied systems, real use

Do these systems hold up outside the lab, in a real operation with real stakes?

Established

Atlas PM built and validated end-to-end from field research inside a real operation; Karyon functional locally, in the builder’s own daily use.

Active

Atlas PM opening its first live deployments; Karyon’s multi-tenant version in development.

Open

Productization and white-label paths — decided by evidence of demand, not by ambition.

Exploring

Longer horizon, clearly speculative — these are directions, not commitments. They happen if the research earns them.

EvaProxy OS — an ACI that is the operating system, not an app running on one.

Dedicated edge neural hardware (NCU) — purpose-built silicon for running composed ACI at the edge.

Follow the research as it happens.