ORIENT RESEARCH · SEPTEMBER 2026 - JAY SOLOMON

ORIENT RESEARCH · SEPTEMBER 2026 - JAY SOLOMON

Owning Organisational Reality

A research programme for machines that can safely maintain organisational reality over time.

A research programme for machines that can safely maintain organisational reality over time.

A research programme for machines that can safely maintain organisational reality over time.

How can a machine safely maintain organisational reality over time?

Organisations do not change when someone asks an AI a question. Decisions move, evidence accumulates, assumptions weaken, people disagree, systems act and yesterday’s valid understanding becomes obsolete. Orient is researching the intelligence required to maintain a persistent model of that changing reality — continuously, safely and economically.

THE RESEARCH THESIS

Judgement is the capability beneath maintained reality.

Orient’s programme studies three operations required to keep organisational understanding coherent as reality moves.

UNDER JUDGEMENT

Form

Form

Transition

Transition

Reconcile

Reconcile

Form what deserves to become state. Transition what evidence changes. Reconcile what accumulated incorrectly.

ACT I · THE PROBLEM

Maintaining organisational reality

ACT II · THE INTELLIGENCE

How Orient judgement is created

ACT III · CONTINUOUS COGNITION

From AI on demand to persistent intelligence

ACT IV · THE PROGRAMME

What we will test and build

ACT I — THE PROBLEM

Maintaining organisational reality

Maintaining organisational reality

The research foundation: how observations become grounded organisational state, how that state changes, and how accumulated errors are found without bypassing judgement.

01 · THE STATE PROBLEM

Search retrieves the record. Orient maintains the current belief.

Monday’s document still exists, remains searchable and may rank highly in retrieval. But organisational reality has changed.

MONDAY · RECORD

Enterprise pricing = €49

The document remains searchable.

TUESDAY · AUTHORITATIVE CHANGE

Enterprise pricing = €65

Leadership changes the price from 1 October.

CURRENT ORGANISATIONAL STATE

€65 from 1 October

€49 remains historically valid until 30 September.

SEARCH ASKS

What information exists?

ORIENT ASKS

What should the organisation currently believe?

02 · THE FUNDAMENTAL ARCHITECTURE

Observation → Interpretation → State

Orient’s training target is not extracted text. It is the appropriate interpretation, epistemic action and successor state.

OBSERVATION

What happened

Immutable organisational history.

messages · meetings · documents · data · agent actions

INTERPRETATION

What does it mean?

Evidence is assessed before state changes.

independence · authority · relevance · supersession

STATE

What should we currently believe?

A grounded, time-aware picture of current understanding.

claims · decisions · expectations · uncertainty · dependencies · safe_to_act

STAGED TECHNICAL FORMULATION

Eₜ = I(Oₜ, Sₜ)

aₜ = πθ(Sₜ, Eₜ)

Sₜ₊₁ = T(Sₜ, Eₜ, aₜ)

New observations are first interpreted as evidence in the context of current state; that interpretation informs an epistemic action; the action then produces—or preserves—the successor state.

What should change—and what must remain?

The same observation can affect a belief, open a question or propose a decision. It does not carry the authority to change every part of the model. Explore two cases from the same starting state.

Starting Understanding

March launch is on track.

Existing Commitment

March promised to the client.

01

Observation

Supplier report: qualification failed.

Retain the source, author and time.

02

Interpretation

Evidence challenges the supplier-readiness assumption.

Interpretation remains revisable.

03

Proposed Change

Revise launch confidence. Flag the March commitment as at risk.

Identify what changes—and what stays.

04

Judgement

Assess evidence and consequence. Escalate any commitment change.

Apply authority and review requirements.

05

Maintained State

March launch is on track.

Existing state · awaiting assessment.

Independent Check

Bounded Reconciliation

Reconstruct the relevant state from evidence without seeing the maintained answer. Compare the two. Any discrepancy becomes a governed correction proposal, not an automatic overwrite.

Continuous Cognition includes justified change and justified preservation.Illustrative research architecture

Reality, intent and continuous judgement.

A maintained model needs to represent both what the organisation understands and what it is trying to accomplish. The research programme treats goals, objectives, principles, commitments and measures as governed state, with owners, scope, authority, effective dates and revision history.

People establish and ratify organisational intent. A machine-inferred objective remains a proposal until an authorised person accepts it. Conflicting goals remain visible; a system must not silently choose a new organisational purpose or rewrite goals to fit an outcome.

Intent informs relevance and consequence. A supplier delay matters differently when it threatens a customer commitment, a safety principle or a launch objective. The judgement question becomes: what changed that matters to what we are trying to achieve?

OrientBench should test whether attention follows current, authorised objectives; whether expired or superseded commitments stop directing work; whether competing goals are surfaced; and whether changes to intent propagate to dependent decisions and agents.

Legibility is a system requirement: people should be able to inspect the evidence, interpretation, proposed transition, authorisation and subsequent outcome. That record explains organisational judgement; it does not claim to expose every internal computation of the underlying model.

03 · THREE MAINTENANCE OPERATIONS

Form. Transition. Reconcile.

FORM

What deserves to become state?

Something being written does not automatically make it organisational truth.

TRANSITION

What should change?

Create, update, reinforce, challenge, supersede, preserve or abstain as evidence arrives.

RECONCILE

What has the model accumulated incorrectly?

Reconstruct bounded areas independently and compare them against maintained interpretation and state.

RECONSTRUCT INTERPRETATION, THEN STATE

Ê⁽ᵏ⁾ = Rₑ(O⁽ᵏ⁾)

Ŝ⁽ᵏ⁾ = Rₛ(Ê⁽ᵏ⁾)

Independent reconstruction catches both false changes and changes Orient missed.

Reconciliation never edits state directly. A discovered discrepancy becomes another proposed transition.

04 · RELIABILITY

A persistent world model can fail in three ways

SPURIOUS MUTATION

It changes something that should remain true.

MISSED MUTATION

Reality changes and Orient fails to follow it.

INCORRECT RESOLUTION

Orient recognises change but moves state to the wrong result.

A system that never changes anything cannot corrupt today’s state — but it will eventually become perfectly accurate about yesterday.

Make the right mutation, at the right time, for the right reason.

EXPECTED STATE LOSS

Reliability must account for every way maintained state can be wrong.

Expected State Loss is a consequence-weighted measure of spurious mutation, missed mutation and incorrect resolution over time.

ACT II — THE INTELLIGENCE

How Orient judgement is created

How Orient judgement is created

The training programme: transform abundant general intelligence into a specialised policy over changes to organisational understanding.

05 · THE CONCEPTUAL BRIDGE

General intelligence already exists

Orient’s research problem is to transform it into specialised organisational judgement.

GENERAL INTELLIGENCE

Broad capability

EPISTEMIC INTELLIGENCE

Evidence, authority, uncertainty and restraint

ORGANISATIONAL STATE INTELLIGENCE

How understanding should change

Orient is not principally training a model that knows more. It is training a policy over changes to organisational understanding.

aₜ = πθ(Sₜ, Eₜ)

An epistemic action policy: given current state and interpreted evidence, choose what—if anything—should change.

06 · THE TRAINING OBJECT

What one Orient training record looks like

The model is not rewarded for extracting a number. It must understand authority, timing, supersession, historical validity and whether action is safe.

STATE BEFORE

Enterprise pricing = €49

Ratified leadership decision

NEW ACTIVITY

“We agreed Enterprise moves to €65 from 1 October.”

Leadership meeting · 3 September

ORIENT JUDGEMENT

SUPERSEDE €49 · CREATE €65 · EFFECTIVE 1 OCT

PRESERVE historical €49

SAFE_TO_ACT = true

NEW ACTIVITY

“I wonder if we could get away with €70.”

Speculation is not a decision.

CORRECT ACTION

RECORD PROPOSAL

CURRENT STATE: UNCHANGED

EPISTEMIC RESTRAINT

Do not mutate the ratified price.

Request evidence or wait for authority.

SAFE_TO_ACT = false

07 · THE COMPOUNDING DATA ASSET

The judgement corpus records how reality should change

Ordinary enterprise AI accumulates documents and interactions. Orient accumulates examples of what new evidence should cause an organisation to understand differently.

01

STATE BEFORE

02

NEW ACTIVITY

03

EVIDENCE

04

PROPOSED ACTION

05

HUMAN / MACHINE JUDGEMENT

06

STATE AFTER

07

LATER RECONCILIATION

The documents are not the deepest proprietary data. The asset is the judgement about what those documents caused the organisation to understand differently.

08 · THE INTELLIGENCE FACTORY

Models are outputs of the system, not the moat themselves

GENERAL / FRONTIER MODELS

teacher judgement · alternatives · uncertainty · hard negatives

ORIENT JUDGEMENT CORPUS

action-labelled history of how reality should change

ORIENT FAST

High-volume perception and obvious cases

ORIENT STATE

Ordinary state judgement

ORIENT JUDGE

Ambiguity, verification and consequence

EVALUATION

ORIENTBENCH

DEPLOY

DIFFICULT CASES · LEARNING PATH → JUDGEMENT CORPUS

DIFFICULT CASES · LEARNING PATH → JUDGEMENT CORPUS

Model weights are outputs of the intelligence-production system, not the moat themselves.

THE INTELLIGENCE-PRODUCTION SYSTEM

The model weights are only one output

Building Orient intelligence requires more than training model weights. It requires the system that produces, evaluates and continuously improves those models.

The programme will therefore build five connected assets.

The model weights are outputs of this system. The system itself is the durable capability.

Five connected assets

Together, these assets turn general intelligence into a repeatable system for organisational state judgement.

THE ORIENT JUDGEMENT CORPUS

STATE BEFORE + NEW ACTIVITY + RELEVANT EVIDENCE + CANDIDATE EPISTEMIC ACTION + JUDGEMENT + STATE AFTER + LATER RECONCILIATION

A proprietary dataset recording not merely what an organisation said, but what new evidence should have caused its understanding to become.

ORIENTBENCH

A dedicated evaluation environment for state formation, mutation precision and recall, spurious and missed mutations, incorrect resolution, temporal reasoning, authority, contradiction, abstention, consequence, calibration, safe-to-act, reconciliation and correlated model failure.

THE ORIENT ROUTER

A learned policy for determining the minimum sufficient intelligence required to reach a safe organisational judgement, escalating when ambiguity or consequence demands it.

THE RECONCILIATION SYSTEM

Mechanisms that independently reconstruct selected areas from underlying observations, compare them with maintained state, and return discrepancies as proposed transitions through the same governed judgement process.

THE INTELLIGENCE-PRODUCTION PIPELINE

A repeatable process combining teacher generation, hard-negative mining, human judgement, preference learning, distillation, calibration, model comparison, routing optimisation and reconciliation feedback.

A model enters production because it improves organisational-state performance — not because it scores well on a generic language-model benchmark.

The result is not one model, but a governed system in which specialised models, evaluation, routing and reconciliation improve together.

The target architecture

ORGANISATIONAL ACTIVITY

↓ ORIENT FAST

↓ ORIENT STATE

↓ ORIENT JUDGE / DEEP WHEN REQUIRED

↓ MAINTAINED ORGANISATIONAL REALITY

↓ SAFE HUMAN AND MACHINE ACTION

↺ RECONCILIATION CONTINUOUSLY TESTS WHETHER THE STATE REMAINS GROUNDED

This creates a transition from general-purpose AI to something much more specialised:

General Intelligence → Epistemic Intelligence → Organisational State Intelligence

The first research milestone is concrete:

Can an Orient-specialised model outperform much larger general-purpose models on the epistemic operations required to maintain organisational reality — while using substantially less computation?

If it can, the programme then asks how far that intelligence can be compressed, specialised and distributed across the model family without increasing Expected State Loss.

That is what the compute is being used to discover.

09 · THE PROMOTION GATE

OrientBench — measuring organisational judgement

Generic benchmark improvement does not matter unless the model maintains organisational reality more accurately, safely and efficiently.

01

Formation

Does valid activity become state?

02

Mutation precision

Were changes justified?

03

Mutation recall

Were genuine changes found?

04

Spurious mutation

Did stable state change incorrectly?

05

Missed mutation

Did obsolete state survive?

06

Incorrect resolution

Did state move to the wrong result?

07

Temporal reasoning

Were effective dates understood?

08

Authority

Was decision power interpreted correctly?

09

Evidence interpretation

Was evidence relevant, independent and sufficient?

10

Contradiction

Were incompatible claims preserved?

11

Abstention

Did Orient refuse unsupported change?

12

Consequence estimation

Consequence uncertainty is distinct from state uncertainty: Orient must estimate both the likely downstream impact of an error and how uncertain that estimate is.

13

Calibration

Did confidence match correctness?

14

Safe-to-act

Could action safely follow state?

15

Reconciliation

Did reconstruction detect drift?

16

Error covariance

Agreement is weak evidence when models share failure modes: correlated errors can make several models confidently wrong together.

A model is not better because it scores higher generally.

A stronger result means lower Expected State Loss at a comparable compute budget, or lower compute cost while meeting the same reliability threshold.

Reliability within a compute budget.

We will compare architectures at matched workloads and compute budgets, measuring spurious mutations, missed mutations and incorrect resolutions over time. Each error is weighted by its downstream consequence; high-consequence failures are reported separately so averages do not conceal them.

The complementary economic test is the compute cost required to meet a defined reliability threshold. Dividing Expected State Loss by compute is not a standalone optimisation objective: spending more must not look like progress unless state quality improves.

Baselines should include capable low-cost general models, frontier models and specialised portfolios with selective escalation. Evaluation must include the cost of routing, verification and reconciliation, as well as the primary inference.

Before reporting cost per correctly maintained unit of organisational reality, OrientBench must define the unit, evaluation horizon and correctness criteria. Candidate units include a maintained claim or decision with its evidence, temporal validity and dependencies.

Actionability is conditional on the proposed action, actor, authority and current evidence. Human ratification remains attributable and revisable; agreement among models must be tested for shared failure modes.

External reference tests include AA-Briefcase for knowledge work, GDP.pdf for document reasoning, AutomationBench-AA for cross-application workflows, and long-context and reliability evaluations. These help select candidates; they do not substitute for Orient-specific state evaluation. Record the Intelligence Index version rather than comparing scores across benchmark revisions.

OrientBench must additionally carry state forward through sequences of new evidence: preserve valid history, recognise supersession, retain unresolved contradictions, identify authority, surface affected decisions and route proposed changes for judgement. The test is whether organisational understanding remains justified through time.

ACT III — CONTINUOUS COGNITION

From AI on demand to persistent intelligence

From AI on demand to persistent intelligence

The runtime thesis: specialised intelligence can continuously maintain organisational state when computation is allocated according to epistemic need.

10 · THE PRODUCT IMPLICATION

AI no longer needs to wait to be asked

Search operates when queried. Agents operate when tasked. A live organisational model must operate continuously.

SEARCH

question → answer

Operates when queried.

AGENT

task → action

Operates when tasked.

ORIENT

activity → judgement → state → continuous reconsideration

Operates continuously.

Continuous Cognition maintains organisational understanding through time and determines what deserves reconsideration, investigation, correction or escalation—even when nobody has asked a question.

11 · CONTINUOUS RECONSIDERATION

A live model operates continuously

New activity does not merely enter a corpus. It can cause previously settled understanding to be examined again.

Evidence arrives

Revisit affected beliefs.

A decision changes

Inspect everything downstream that depended upon it.

Contradictory evidence arrives

Reopen previously resolved understanding.

Evidence ages

Reconsider confidence.

An agent wants to act

Recalculate whether the underlying state remains safe_to_act.

What deserves cognition next?

Continuous cognition means continuous eligibility for attention, selectively exercised. A change in evidence, approaching deadline, ageing belief or unresolved contradiction can trigger reconsideration without a new human prompt. Most state receives no computation most of the time.

Attention allocation selects which claims, questions or situations deserve examination now, based on change, uncertainty, dependencies, consequence and relevance to the organisation’s current authorised goals and commitments.

Model routing selects the minimum sufficient model or reasoning process for that examination. The policy escalates when uncertainty, potential consequence or uncertainty about that consequence warrants it.

Reconciliation scheduling selects bounded areas of accumulated understanding for independent reconstruction. Reconstruction is blind to the maintained answer; discrepancies return as proposed transitions rather than silently overwriting state.

The programme will evaluate attention, routing and reconciliation together: did the system examine the right state, spend enough computation, and find consequential drift in time?

12 · CONTINUOUS COGNITION ECONOMICS

Why specialised models make continuous cognition possible

Continuous cognition changes the economics of organisational AI.

If every message, document, meeting and system event had to be analysed by frontier intelligence, continuously maintaining organisational state would be unnecessarily expensive.

Orient is designed differently.

Most organisational activity requires relatively narrow forms of judgement: identifying relevant evidence, linking entities, recognising decisions, detecting straightforward changes, or determining that nothing material has changed. These operations can be performed by smaller models specialised specifically for organisational state.

More difficult cases are routed upward.

ORIENT FAST

Orient Fast can operate continuously across high volumes of activity.

ORIENT STATE

Orient State handles meaningful state transitions.

ORIENT JUDGE

Orient Judge examines uncertainty, consequence and contested changes.

ORIENT DEEP

Orient Deep or frontier intelligence is reserved for the relatively small number of cases where substantially deeper reasoning is justified.

The target is a different cost structure:

cheap intelligence continuously, stronger intelligence selectively, frontier intelligence rarely.

This is why building specialised Orient models matters.

The goal is not merely to reduce the cost of individual model calls. It is to reduce the cost of maintaining organisational reality far enough that intelligence can operate continuously.

Orient can therefore revisit affected claims when evidence arrives, reconsider contradictions, recalculate confidence, inspect downstream dependencies, test safe_to_act, and reconcile accumulated state even when no human has asked it to do so.

THE RELEVANT ECONOMIC UNIT

The relevant economic unit is therefore not cost per token or cost per model call. It is:

Cost per correctly maintained unit of organisational reality

THE QUESTION IS NOT

What does one token or model call cost?

THE QUESTION IS

What is the minimum compute required to make this epistemic judgement reliably?

Compute follows risk-adjusted epistemic need

Allocate attention → select sufficient intelligence → evaluate state quality and cost

UNCERTAINTY

How ambiguous is the judgement?

CONSEQUENCE

What happens if it is wrong — and how uncertain is that impact?

CAPABILITY

Which model can resolve it reliably?

COST

What is the minimum sufficient computation?

The research hypothesis is that specialisation can make routine judgement inexpensive, routing can reserve deeper reasoning for the cases that need it, and reconciliation can detect accumulated errors. The combined reliability and cost must be measured.

13 · THE RUNTIME HIERARCHY

Minimum sufficient intelligence, escalated by need

The exact model sizes are research outputs, not fixed product requirements.

↓ DECREASING EVENT VOLUME

↓ DECREASING EVENT VOLUME

INCREASING UNCERTAINTY / CONSEQUENCE ↑

INCREASING UNCERTAINTY / CONSEQUENCE ↑

01 · ORIENT FAST

High-volume perception, relevance and obvious cases

02 · ORIENT STATE

Ordinary organisational state judgement

03 · ORIENT JUDGE

Ambiguity, verification and consequence

04 · ORIENT DEEP

Difficult reconciliation and long reasoning

05 · HUMAN

Irreducible high-consequence ambiguity

FUTURE RESEARCH FRONTIER

Harder cases may need more computation, not a permanently larger model

SIMPLE OBSERVATION

one pass

ORDINARY MUTATION

standard reasoning

CONTRADICTION

deeper / recurrent reasoning

DEEP AMBIGUITY

more passes

HIGH CONSEQUENCE

human

ACT IV — THE PROGRAMME

What we will test and build

What we will test and build

OrientBench, six research workstreams, Arrhenius-scale experimentation and the compounding intelligence-production system they are designed to create.

THE RESEARCH OUTPUT

Orient organisational state intelligence

The programme is intended to produce a new class of specialised AI: Orient organisational state intelligence.

The objective is not to train another general-purpose foundation model from scratch.

General intelligence already exists.

Orient will use capable open and frontier models as starting points, teachers and judges, then specialise that intelligence around a much narrower problem.

The intelligence is specialised around one question:

Given what an organisation currently understands, what it is trying to achieve and what just happened, what—if anything—should now change?

The first concrete output

The programme targets an Orient model portfolio specialised for maintaining organisational state. External models supply starting points, teachers and comparison baselines; the final role allocation is determined by evaluation.

Rather than assuming that one model should perform every operation, the research will determine the smallest and most capable intelligence required for each role.

ORIENT FAST

High-throughput perception and triage: candidate entities, claims, evidence, decisions, temporal signals and relevance. Tiny workers such as MiniCPM5-2B are candidates for these constrained operations, not default arbiters of organisational truth.

ORIENT STATE

Governed state maintenance: determine whether evidence should create, reinforce, challenge, supersede or preserve understanding. Current experiments prioritise GLM-5.3-Flash and Qwen3.8-Flash-Next, with other efficient models as comparators.

ORIENT JUDGE

Verification and escalation: assess evidence, authority, time, consequence, uncertainty and actionability. Kimi K3 and GLM-5.3 are primary teacher/judge candidates; independent judgement must be tested for shared failure modes.

ORIENT DEEP

The strongest reasoning layer for difficult contradictions, long temporal situations, unusual ambiguity, high-consequence changes and complex reconciliation. It may use a larger specialised model, adaptive or recurrent computation, frontier intelligence, or a combination of them.

These are research roles, not four fixed checkpoints. Model selection, specialisation, deployment size and routing thresholds remain outputs of the programme.

What matters is the resulting capability.

The programme may discover that some roles can be combined, that particular operations benefit from separate specialists, or that additional computation within a smaller model performs better than invoking a larger one.

Smiling woman in a dark blazer against a blue sky
Red table lamp on a pink stone counter in a wood-panelled interior
Two people talking at a table beside an open laptop

A model portfolio, selected by organisational role.

RESEARCH SHORTLIST · 9 SEPTEMBER 2026

Orient Fast, State, Judge and Deep name the capabilities we are researching. The external models below are candidates, teachers and baselines—not announced production deployments or renamed Orient models. The shortlist changes as evidence improves; OrientBench determines promotion.

Tiny workers · Orient Fast

MiniCPM5-2B joins the high-volume worker experiments for candidate entities, claims, classification and routing signals. It must demonstrate schema reliability and appropriate abstention before promotion. Final contradiction resolution and state adjudication remain separate responsibilities.

Continuous state · Orient State

GLM-5.3-Flash and Qwen3.8-Flash-Next lead the current experiment shortlist. DeepSeek V4 Flash and Mistral Small 4 remain comparators. Evaluate evidence attribution, temporal updates, contradiction detection, multimodal interpretation and governed low-risk state changes—not general benchmark rank alone.

Teacher and judge · Orient Judge / Deep

Kimi K3 and GLM-5.3 are the primary open-weight teacher and adjudication candidates. Qwen3.8-2.4T-A95B remains a secondary long-context text comparator, rather than a primary judge candidate. Its downloadable checkpoint must not be confused with the additional capabilities of hosted Qwen3.8-Max. Model agreement must be evaluated for correlated failures.

Multimodal retrieval · evidence access

WeMM-Embedding-2B joins the retrieval shortlist alongside Qwen3-VL-Embedding-2B and Jina v5 Omni. Test retrieval across emails, documents, screenshots, charts, slides and UI captures. Vendor-reported benchmark gains are hypotheses to reproduce on organisational evidence. Retrieval supplies candidate evidence; it does not decide what the organisation should believe.

Open research · reproducible experiments

K2 Horizon remains an important reproducibility and long-context research candidate. Its role is to support transparent experimentation, comparison and specialisation; it is not presumed to replace the primary teacher/judge candidates.

15 · THE RESEARCH PROGRAMME

Six workstreams connect the science, system and economics

WP1

FORMATION & INTERPRETATION

Can machines reliably determine what organisational activity deserves to become evidence and state?

WP2

STATE TRANSITION

Can they update, preserve, contest and abstain correctly?

WP3

CONSEQUENCE & SAFE ACTION

Can they determine what a mutation could affect and whether downstream action is safe?

WP4

RECONCILIATION & LONG-HORIZON INTEGRITY

Can independent reconstruction detect interpretation drift, false mutations and missed mutations?

WP5

SPECIALISATION, PORTFOLIO & ADAPTIVE COMPUTE

Which combination of specialised models, routing and variable reasoning depth minimises Expected State Loss within a compute budget—or minimises compute while meeting a defined reliability threshold?

WP6

JUDGEMENT LEARNING

Can action-labelled judgement data, hard cases, ratification and reconciliation create specialised intelligence that outperforms generic models?

THE OPPORTUNITY

From European compute to European AI capability

Europe is making an enormous investment in sovereign AI infrastructure. The important question is what that infrastructure ultimately creates.

Orient offers a concrete test.

Some of the intelligence required to maintain Orient’s Organisational State Model is currently rented from proprietary frontier-model providers. We propose using European supercomputing infrastructure and the rapidly improving open-weight ecosystem to determine how much of that critical capability can instead be adapted, trained, evaluated and ultimately owned here.

The objective is not another general-purpose frontier model, nor to compete with US and Chinese labs at everything they do. It is to turn increasingly commoditised general intelligence into specialised European capability for problems that matter to European industry.

For Orient, that problem is maintaining an Organisational State Model: a continuously updated representation of what the organisation currently believes, what evidence supports it, what has been decided, what has been superseded, what remains contested, and where uncertainty remains.

For Orient, that means specialised AI capable of forming and maintaining grounded organisational state as the organisation changes.

What this programme would demonstrate

European compute can create proprietary European AI capability, rather than simply providing cheaper access to foreign AI.

Open models can be turned into specialised industrial intelligence, trained around European products, requirements and use cases.

European companies can reduce dependence on proprietary frontier APIs while continuing to benefit from frontier-model progress.

Critical AI capability can run on European-controlled infrastructure, with greater control over deployment, data, models and economics.

Public compute can become private-sector leverage. Infrastructure funded at European scale can allow startups to undertake R&D programmes that would otherwise require substantial private capital.

Europe does not have to win only by building the world’s largest general model. It can also win by becoming exceptionally good at turning increasingly commoditised general intelligence into specialised, defensible systems for industry.

Strategic control, not isolation

AI sovereignty cannot mean isolating Europe from the global model ecosystem. Nor does using an open model originally developed outside Europe suddenly make the entire stack European.

The practical objective is strategic control: the ability to choose the underlying model, run it on infrastructure we control, adapt it ourselves, evaluate it against our requirements, retain our data and judgement layer, replace the model when something better appears, and own the specialised capability created above it.

That is the capability this programme is designed to test.

For Orient, success means increasingly owning the intelligence-production system that keeps organisational reality coherent: its state model, judgement corpus, evaluation, specialised models, routing and learning loop.

For Sweden AI Factory and EuroHPC, success would demonstrate the full chain that sovereign compute is intended to enable:

European infrastructure

→ access for European startups

→ model experimentation and post-training

→ proprietary datasets and evaluation

→ specialised European AI capability

→ deployable European products

→ stronger European technological independence.

The question is therefore larger than whether Orient can reduce its model bill.

Can European public compute give a European startup the ability to take frontier intelligence, specialise it around an important industrial problem, and turn it into a capability that Europe owns and can deploy on its own terms?

If the answer is yes, that is a proof point not only for Orient, but for the European AI Factory model itself.

Large compute to discover intelligence. Small compute to deploy it.

14 · ARRHENIUS

Large compute to discover intelligence. Small compute to deploy it.

Arrhenius is not the infrastructure required to run every Orient customer. It is the factory in which specialised organisational intelligence is discovered.

382

382

GPU NODES

4

4

GRACE HOPPER / NODE

29 PB

29 PB

STORAGE

ARRHENIUS COMPUTE

↙ ↓ ↓ ↘

MODEL TOURNAMENTS

MODEL TOURNAMENTS

same benchmark, many candidates

FINE-TUNING SWEEPS

FINE-TUNING SWEEPS

specialisation at scale

SIMULATED ORGANISATIONS

SIMULATED ORGANISATIONS

controlled long-horizon worlds

LONG-HORIZON EVALUATION

LONG-HORIZON EVALUATION

state drift over 50K events

WHAT THE COMPUTE ACTUALLY DOES

Search a wide experimental space, then deploy only what works

TEACHER GENERATION

SYNTHETIC TRANSITIONS

HARD-NEGATIVE GENERATION

SUPERVISED TRAINING

PREFERENCE TRAINING

DISTILLATION

CALIBRATION

ARCHITECTURE SWEEPS

RECONCILIATION EXPERIMENTS

ADAPTIVE-DEPTH EXPERIMENTS

ORIENTBENCH EVALUATION

ROUTING-POLICY TRAINING

EUROPEAN CAPABILITY PATH

European infrastructure → startup access → model experimentation → proprietary judgement data and evaluation → specialised European capability → deployable products

European compute becomes the enabler of the programme—not the opening reason for it.

16 · THE COMPOUNDING SYSTEM

What Orient ultimately owns

Owning specialised organisational intelligence means owning the system that continuously produces and evaluates it.

ORGANISATIONAL STATE MODEL

JUDGEMENT CORPUS

EPISTEMIC ACTION SPACE

ORIENTBENCH

MODEL PORTFOLIO

ROUTING POLICY

RECONCILIATION

TRAINING PIPELINE

THE RESULT

ORIENT INTELLIGENCE

The weights are replaceable.

The intelligence-production system compounds.

CONCLUSION

From AI on demand to persistent organisational intelligence

From AI on demand to persistent organisational intelligence

General intelligence is becoming abundant.

Orient’s research asks how that intelligence can be transformed into specialised judgement that continuously maintains what an organisation believes, why it believes it, what has changed, and what humans and machines are safe to do.

The system can afford to think when nobody has asked it a question.

That is continuous cognition.

A

Sources and current model references

NAISS, Arrhenius resource overview

https://www.naiss.se/resource/arrhenius/

Sweden AI Factory, August 2026 overview

Startup support, compute access, EuroHPC application assistance, training, expertise, and publicly funded services.

Z.ai, GLM-5.3 announcement, 14 August 2026

https://z.ai/blog/glm-5.3

Z.ai / Hugging Face, GLM-5.3 weights and model card

https://huggingface.co/zai-org/GLM-5.3

Z.ai / Hugging Face, GLM-5.3 License

https://huggingface.co/zai-org/GLM-5.3/blob/main/LICENSE

Z.ai / Hugging Face, GLM-5.3-Flash

https://huggingface.co/zai-org/GLM-5.3-Flash

Artificial Analysis, model leaderboard and comparisons

https://artificialanalysis.ai/leaderboards/models

Qwen, Qwen3.8-27B model card

https://huggingface.co/Qwen/Qwen3.8-27B

Qwen, Qwen3.8-Flash-Next model and licence

https://huggingface.co/Qwen/Qwen3.8-Flash-Next

Shortlist reviewed 9 September 2026. Record checkpoint, deployment mode, benchmark version, licence and evidence date for every comparison. Benchmark revisions can change rankings without any change to model weights. Active parameters alone do not determine hardware footprint or serving cost.