Model Watch
MODEL WATCH · UPDATED 21 SEPTEMBER 2026
Research question → candidate → evaluation → evidence status → product implication. This dated view distinguishes current user-experience integrations from research candidates. Deployment does not establish research performance; vendor-reported claims and results reproduced by Orient must be labelled separately.
EVIDENCE STATUS · The entries below describe integrations and proposed evaluations. They do not report completed comparative OrientBench results. Current use should not be read as proof of suitability for live state maintenance.
USED IN THE USER EXPERIENCE
Fast typed judgement · Jev
What can fast typed judgement contribute to a responsive user experience?
Jev currently supports intent classification and routing in Ask and Think, and semantic filtering in Activity and Events. It is not yet used in the live organisational state model. Broader state-maintenance evaluation is planned research, and will require separate evidence.
WATCHLIST · NOT A PRODUCTION REPLACEMENT
Open evaluators · OpenJev approaches
Can bounded organisational judgements be trained and deployed on infrastructure we control?
Compare Dasein’s local option scorer and task-specific training approach with AlexWortega’s natural-language inference model. Record the exact checkpoint, licence, task shape and calibration. API similarity or a shared name does not establish Jev-equivalent behaviour.
Compact generative workers · experiment shortlist
MiniCPM5-2B joins the high-volume worker experiments for candidate entities, claims, classification and routing signals. It must demonstrate schema reliability and appropriate abstention before promotion. Final contradiction resolution and state adjudication remain separate responsibilities.
State maintenance · experiment shortlist
GLM-5.3-Flash and Qwen3.8-Flash-Next lead the current experiment shortlist. DeepSeek V4 Flash and Mistral Small 4 remain comparators. Evaluate evidence attribution, temporal updates, contradiction detection, multimodal interpretation and governed low-risk state changes—not general benchmark rank alone.
Teachers and judges · experiment shortlist
Kimi K3 and GLM-5.3 are the primary open-weight teacher and adjudication candidates. Qwen3.8-2.4T-A95B remains a secondary long-context text comparator, rather than a primary judge candidate. Its downloadable checkpoint must not be confused with the additional capabilities of hosted Qwen3.8-Max. Model agreement must be evaluated for correlated failures.
Multimodal retrieval · experiment shortlist
WeMM-Embedding-2B joins the retrieval shortlist alongside Qwen3-VL-Embedding-2B and Jina v5 Omni. Test retrieval across emails, documents, screenshots, charts, slides and UI captures. Vendor-reported benchmark gains are hypotheses to reproduce on organisational evidence. Retrieval supplies candidate evidence; it does not decide what the organisation should believe.
Reproducibility · experiment shortlist
K2 Horizon remains an important reproducibility and long-context research candidate. Its role is to support transparent experimentation, comparison and specialisation; it is not presumed to replace the primary teacher/judge candidates.
WATCHLIST · NEW CANDIDATE
Local workers · Xing4.0
Which interpretation and tool-use tasks can a local generative worker handle reliably?
Xing4.0-29B-A4B is a candidate for local worker comparisons. Test evidence attribution, proposal-versus-decision distinctions, valid tool arguments and abstention. Measure the full memory and serving requirements; active parameter count alone is insufficient.
WATCHLIST · DEPLOYMENT EXPERIMENT
Compression · Ternary Bonsai 2
What happens to state-maintenance accuracy under tighter memory and compute budgets?
Bonsai 2 is a compression candidate. Compare it with the underlying model and conventional quantisation on the same organisational tasks. Reproduce vendor performance claims on the intended runtime, and measure working memory, context cost and consequential errors. Compression and distillation are separate experiments.
RESEARCH TRACK · CANDIDATE SELECTION ONGOING
Multimodal interpretation · original evidence
Can original audio, documents and visual evidence preserve distinctions lost in text extraction?
Compare native multimodal interpretation with transcription or document extraction followed by text reasoning. Test speaker attribution, qualifications, disagreement, charts and the distinction between a proposal and a decision. Separate hosted API capabilities from what an available open checkpoint supports.
For every comparison, record the checkpoint or API version, access and licence, serving configuration, evidence date, benchmark version and measured workload. Label vendor-reported results separately from reproduced Orient results.